# STATUS — what works, what's broken, what's next **Ready for the first real tester (Tester-2): yes. Tester 2 is a laptop that is switched off at night (your word, 2026-10-05) — it was offline all session; nothing was sent to it.** **Updated 2026-10-05 (day, the night's fixes): every box of ours healthy. Fixed and proven live: the off-site clean-up now deletes old copies (both demo boxes), a new box's first app install, the update's disk-space check, a killed update's lost report. The power cut in the middle of an update was tested on demo-hp with your go: the box came back by itself in 37 s, but the next update failed until I ran one command by hand — filed (R-876), fix next session. One decision for you below (a box that is off at night). Report: `REPORT-night-fixes-2026-10-05.md`.** ## Today (2026-10-05, day): the night's fixes, the power cut by day, a box that is off at night **Decisions I took myself (you may reverse each — `09` decisions 104–108):** - The off-site clean-up's safety line is now "the last 7 calendar days" — the same number the clean-up keeps — not "8 days old". It can never block an honest clean-up again. - The image clean-up now waits while ANY app install, update, restore or undo is downloading — not only installs. - A killed update keeps its report on the box until the hub has it; the agent looks for such reports every 5 minutes. - **A second agent release today (0.144.1)**, against "one release per repo": the first fix for the lost report was proven NOT to work on demo-hp, and shipping it as it was would have been worse. **What works now (proven live):** - **The weekly off-site clean-up really deletes old backups.** One clean-up each by hand: demo-felhom 16 → 14, demo-hp 145 → 127 — exactly the backups I predicted. No error mail, the hub's count check quiet, the key files clean. This also closes R-95 (the box can no longer delete its own off-site history, and clean-up now works). - **A new box's first app install works the first time:** the image clean-up met an install at minute 3 on the scratch box, waited, and BookStack installed first try. A failed install now logs its real reason. - **The update's disk-space check counts the real download** (12.8 MB for 13 packages; it counted 0 before). - **A killed update still reports to the hub** (5 minutes later, once). The debug update runs with the hub away. - Controller 0.294.0 on demo-hp, demo-felhom and Tester 1 (floor per customer; Tester 2 not moved). Agent 0.144.1 and its root files on all three. New-install image 0.294.0 baked and approved. **Found today (filed, not fixed):** - **After a power cut in the middle of an update, every later update fails until someone runs one command on the box** (R-876, P2). The box itself comes back fine. Until the fix: `runbooks/crash-guard.md` has the command. I fix it next. - A box that is off at night (below): no catch-up, no missed-backup alarm (R-872), a "server cannot be reached" mail to the household every night (R-873), restore-tests never run (R-874). **Needs you:** 1. **A box that is off every night (Tester 2) — what does the product promise?** Today such a box never gets its nightly database backups, second copy or off-site copy; the whole-box backup runs only about every 2 days; no alarm says so; the household is mailed "your server cannot be reached" every night. No design document covers it (R-871). - **A (my pick): a missed night runs once when the box comes back.** The database backups, the second copy and the off-site copy run a few minutes after the box is on again (they take seconds to minutes; whether any of them pauses an app is to be measured in the design — the whole-box backup, which does pause apps, already has its own catch-up). App and system updates still wait for a night. Costs: a design section and one controller release; the household may notice a busy disk for a few minutes after switching on. - **B: say plainly that the box must stay on at night.** The setup guide and the box's backup page say it; the missed-night alarm fires after 2 nights off. Costs: wording + one alarm; a laptop household gets an alarm it cannot fix except by changing habits. - **If you do nothing:** Tester 2 keeps having no database or off-site backup, nobody is told, and the household keeps getting the nightly "cannot be reached" mail. A restore-test alarm will fire around 2026-10-11. 2. **Tester 2's one-time step** — unchanged from yesterday (below). If you wait, it keeps working; it just cannot get new root files. **Recorded, not decisions:** Tester 1's Cloudflare tokens are NOT rotated (your ruling; R-870 has the steps). ## Today (2026-10-04, night): root files for installed boxes, test approvals, Docker self-repair **Decisions I took myself (you may reverse each):** - The hub cancelled only the AUTOMATIC test approvals of today (guest and host fixes). Your own "Approve Docker set" press stays in force. From now on every approval made during a test wait gets the test mark, button or not. - A box whose root files are behind the approved ones for 7 days sends you a mail. - The demo boxes' controller floor is 0.293.0; the fleet floor stays 0.292.0, so Tester 2's controller did not move. **What I did:** - **Root files for installed boxes (your ruling):** a box's root-owned files (permissions list, helper scripts, crash guard) now travel as one signed package. The box checks your signature, every file and itself; on any problem it puts the old files back. Proven on both demo boxes: a wrong package refused, a package with one changed line installed and undone, a copied old job refused. New installs use the same package. - **One catch:** a box installed before tonight cannot take the FIRST package by itself; it needs one small step by hand. I did it on both demo boxes. **Tester 2 needs it from you** (below). - **Tester 2 was not as old as the brief thought:** it was installed at 18:06 local, after the crash guard and the new image. It already has the crash guard, live-restore and Docker 29.8.2. It lacks only tonight's Docker-update fix. I sent it the signed agent update. - **Test approvals end with the test:** the hub cancelled today's 4 test approvals at its restart (you got one mail). Tester 2 keeps what it installed; no further box installs them. The same fixes get a real approval after 24 h + a night. - **Docker self-repair:** if Docker's socket is re-created, the controller now restarts itself and the web router within about 2 minutes. Proven on the scratch box and on demo-hp, no app restarted. - **drill-r50 is gone** from the hub. On ep0 nothing was destroyed (it had no backups there); its tunnel entry left. - The "felhom-pbs skipped" line on Tester 2 is normal for a new box's first hour (no mail was sent). - New-install image 0.293.0 baked and approved. **Needs you (nothing breaks if you wait):** - **Tester 2 is offline** since 20:06 local (and was off 19:13–20:05 local; it restarted in between). I sent it nothing before 20:35. My agent update for it is queued but expires at 21:20 local; when the box is back I re-send it. Worth asking the tester whether the box was switched off. - **Tester 2's one-time step** (5 minutes): connect your tunnel, `ssh -p 8822 felhom-op@10.77.0.5`, reveal the root password in the hub (Hosts → Tester-2 → Console access; this writes one line on Tester 2's timeline), `su -`, then run the three commands in `documentation/runbooks/config-bundle.md` ("Tester 2"). Tell me when done; I send the package. If you wait: Tester 2 keeps working; it just cannot get new root files until then. - **Found, not fixed:** the agent's permission list is wider than "minimal": a broken-into agent could become root on its own box. Worth fixing before the first paying customer. ## Today (2026-10-04, ~18:30): you found a bug — the Docker update blinded the box's controller - **What happened:** the Docker update on the N100 restarted Docker itself. The apps kept running (as designed), but the controller and the web router kept a connection to the OLD Docker, so the controller could not see anything: the hub showed the N100 DOWN from 14:18 to 15:57. demo-hp had the same fault; the crash test happened to heal it. - **Fixed (your choice):** after a Docker update the box now restarts just those two (about 10 seconds, apps untouched), and the update's health check now asks "can the controller really reach Docker?". Released as host agent 0.142.1, proven twice on demo-hp, on both demo boxes now, and approved for new installs. ## Today (2026-10-04, late evening): the System page, Docker updates, the crash restart **Decisions I took myself (you may reverse each):** - The box knows a crash only as "it did not shut down cleanly" (the crash memory chip saved nothing). So a power cut also counts as a crash. - An "oops" (a kernel error the box survives) does not restart the box; you get a mail instead. - Your words win where the brief disagreed: the **3rd** crash within one hour leaves the box off. Restart after 10 s, re-arm after 24 h. All settings. - Only the box itself decides whether a Docker update is allowed: it checks your signature with a key file only root can change (not the agent's own settings, which the agent could change). - The System page uses the alarm limits for its colours (red = an alarm would fire). **No decision needed from you today.** **What I did:** - **The System tab** in the hub: per box the Proxmox, kernel (now and next boot), Debian and Docker versions, what is waiting, held packages, "restart needed", the crash guard and the last update run — with buttons for ring, on/off, "Approve now" and "Approve Docker set". The Hosts page shows Proxmox and kernel too. - **Docker updates:** "live-restore" is on in every box (no app restarted: 24 of 24 and 5 of 5 containers kept running). Both demo boxes moved to Docker 29.8.2, every app kept running. I approved that set with the new button (a 0-night test wait, then back to 2 nights). An undo signed by your key put demo-hp back one version and forward again, apps running throughout. A copied (replayed) signed job was refused. - **Crash restart (your 3 crashes on demo-hp):** crash 1 and 2 — back by itself in under a minute; crash 3 — it stayed off until you switched it on. You got the "guard tripped" mail. I re-armed it. - **Found:** the hub learns about a crash up to 15 minutes late (nothing lost). After a crash, the household can also get an "app stopped" mail besides "restarted after a crash" — whether to calm that is a later choice for you. - **New-install image re-made** with live-restore on and the approved Docker version, and approved in the hub with host agent 0.142.0. New boxes also get the crash guard (installer 1.30.0). - **Rows:** 6 closed, 1 opened-and-closed the same day, 5 opened. The list went from 334 to 333. **Needs you later (nothing breaks if you wait):** - Installed boxes still get new root files only by hand (the long-standing gap; there are no other boxes today). - Kernel updates are still not built (a hung new kernel would stay — needs a fix first). ## Today (2026-10-04, evening): host fixes, the fleet view, a true tunnel status **Decisions I took myself (you may reverse each):** - Only a box the installer itself recorded as "appliance" gets host updates (a record the agent cannot change). - The four alarm times: no update run for 7 days; "reboot needed" for 14 days; the demo boxes approve nothing for 7 days; a box has fixes nobody approved for 14 days. All four are settings. - I released the host agent and the hub a second time today. The first release said "reboot needed" wrongly; the new 14-day alarm would then have mailed you about boxes you had already rebooted. **One decision for you (a safe default if you say nothing) — the Docker engine updates (designed, not built):** 1. **Turn on Docker's "live-restore" on every box.** With it, a Docker engine update restarts no app (measured: 0 restarts). Without it, every app stops for about 30 seconds per engine update. - **A (my pick):** turn it on — in the new-install image and once on existing boxes. It goes on without restarting anything. Cost: it must never be turned off by a plain restart again (that stops every app and starts none). - **B:** leave it off. Every Docker update then means ~30 seconds of every app being down, at night. - **If you say nothing:** nothing changes; Docker updates stay unbuilt. **What I did:** - **The host's Debian fixes now install themselves**, after the guest's, on the same night run, only on appliances, never a kernel or boot package, never a reboot. Proven: demo-felhom installed 108 host fixes and stayed healthy; all 108 came from Debian. A box made "ring 1" for a test installed exactly the one approved host version it lacked. - **A way to put one host package back by hand** is written and proven on demo-hp (and the test taught it two fixes). - **The fleet view** in the hub: one line per box with its updates, "reboot needed since", and the tunnel. - **The tunnel status is now true:** running, not running, or unknown. I blocked demo-hp's tunnel: after two reports (about 30 minutes) you got a "tunnel down" mail; when I unblocked it, "recovered". A simply stopped tunnel heals itself within 5 minutes, before the hub can even see it. - **The night run is fast:** 23–32 seconds when there is nothing to install (target was under 60). - **The kernel test on demo-hp (your two reboots):** Secure Boot works with it, but GRUB's "boot once" does not work on our boxes: the second plain reboot came up on the NEW kernel again. So a new kernel that hangs would stay. That must be solved before kernel updates. demo-hp now runs the newer kernel, healthy. A watchdog chip exists on demo-hp. - **Found and fixed during the live test:** "reboot needed" was wrong in two ways (fixed in the second release). - **New-install image baked and vouched** (0.292.0 with host agent 0.141.1). The golden waiver is gone: not needed. - **Rows:** 2 closed, 2 opened-and-closed the same day, 3 opened. The list went from 332 to 333. **Needs you later (nothing breaks if you wait):** - **A host that crashes does not restart by itself** (Linux's "panic" setting is off). Changing it changes how every box behaves, so it is your call, together with kernel updates. ## Today (2026-10-04, afternoon): the guest's security fixes install themselves **One decision for you (a safe default if you say nothing):** 1. **Undo for a guest update that goes wrong.** I tested it first, as you asked: Proxmox cannot take a snapshot of a customer box at all (the box is linked to the household's drives, and Proxmox refuses). So there is no automatic undo. Today a failed check stops, mails you, and last night's whole-box backup (minutes old) is the undo, by hand. - **A (my pick):** keep it so. Nothing new to build. A restore takes about 1–3 minutes plus losing what apps wrote since the backup. - **B:** build our own disk snapshot under Proxmox. Automatic, but a new mechanism nobody has tested, with a risk to the disk pool. - **If you say nothing:** A stays. **What I did:** - **Your two choices are built.** A returning household's new box sets the old off-site copy aside on night one and starts a new one (nothing deleted). An approved version that Debian already replaced comes from Debian's dated archive. - **Guest security fixes now install themselves.** Each night, after the whole-box backup, the demo boxes install Debian's fixes. When both demo boxes run the same versions healthy for 24 hours and one night, the hub approves that set; every other box then installs exactly those versions. Docker, the host and the kernel are not touched. - Proven live: each demo box installed 53 fixes and stayed healthy. With a 2-minute test wait the hub approved the set; then I put back 24 hours. demo-felhom, made "ring 1" for the test, installed exactly the 3 approved versions it lacked and nothing newer. A run where I stopped an app failed its check and mailed you (you got that mail). - You can switch OS updates off per box in the hub. They are ON by default. The household sees one line in its timeline: "System security fixes installed". - **The tunnel and two other built-in programs are current** (cloudflared was 4 months old). A new monthly check catches them falling behind. The tunnel came back by itself on both demo boxes within 20 seconds. - **Three bugs found and fixed during the live test**, before release. - **Two gaps found, now rows:** existing boxes cannot receive the new update tool through the product (only new installs, or by hand — the demo boxes got it by hand); and the hub's "tunnel status" for every box always reads "inactive" because it checks the wrong place. - **Rows:** 4 closed (one opened and closed the same day), 5 opened. The list went from 331 to 333. ## Today (2026-10-04, day): off-site closed, operating-system updates measured **Two decisions for you — each has a safe default if you say nothing:** 1. **A returning household's first night.** A new box for someone who had a box before makes no off-site copy, because the old copy was made with a key the new box does not have. - **A (my pick):** the new box sets the old copy aside by itself on its first night, then starts a new one. An un-claimed box already does exactly this. Nothing is deleted; you can put the old copy back. Cost: a small box change. A household that wanted to continue the old copy with its recovery code must do that before night one. - **B:** ask the household on the recovery-code evening. Cost: new screens in two languages. If they skip it, the gap stays. - **If you say nothing:** such a household has no off-site copy until someone presses "start a new off-site backup", and you get a mail each time. Nobody is blocked today (no returning household is waiting). 2. **Approved OS updates, when Debian has already replaced the version.** Debian keeps only two versions of a package. - **A (my pick):** the box then fetches the exact approved version from Debian's own dated archive (snapshot.debian.org, signed by Debian). Measured: 2–3 seconds. Cost: one more outside service we rely on — but in 3 months it would have been needed **zero** times for the packages a box has. - **B:** the box waits for the next approval. Cost: no new service; a box can stay unpatched for one cycle. - **If you say nothing:** nothing is blocked now; the first build step can start with B and switch later. **What I did:** - **A restored box is safe by default.** The automatic restore-test was already safe (measured). The agent's disaster restore now refuses to run next to a live original. Restores by hand use a new safe script: no start on boot, no link to the real drives, network off. All proven live on demo-hp. New agent 0.139.0 on both demo boxes. - **The weekly clean-up cannot get stuck any more.** You can allow one bigger clean-up for one box (hub 0.129.0). Proven on a lab copy: 98 backups, the normal limit refused, the bigger one cleaned to 13, and the next week was normal again. - **A dated check for 12 October** looks at the first clean-up that really deletes old backups. - **OS updates, measured on the demo boxes (nothing built for customers):** - The boxes are far behind: 188 updates per host, about 59 per guest. Every guest runs a different Docker version. - A Debian update of the guest took 24 seconds. No app stopped. The undo (restore the guest) took 73 seconds. - A Debian update of demo-hp's host took 60 seconds. The guests kept running. - A broken (killed) update does not fix itself. Two repair commands fix it in 5 seconds. - A Docker update restarts every app (about 30 seconds silent). A Docker setting ("live-restore") avoids that completely — but switching it off again stopped every app and started none. Recorded as a trap. - The new kernel booted fine, and the box fell back to the old kernel on the next restart. **But** if a new kernel hangs, the box keeps trying it. Someone must then switch it off and on. No hardware watchdog is used. - About 40 packages from Proxmox have ordinary names (for example the disk system ZFS and the Secure Boot loader). So "which lane" must follow where a package comes from, not its name. - The internet tunnel program on every box (cloudflared) is 4 months old. Nothing updates it. New row. - **Thank you for being near the box.** demo-hp restarted twice. It now runs the old kernel and starts the new one at its next restart. - **Rows:** 2 closed, 5 opened. The list went from 328 to 331. ## Today (2026-10-04): off-site safety finished - **The weekly clean-up is ON and no longer stops itself.** The fake-backup check now skips young copies that a manual backup replaced the same day, instead of refusing. I ran one window on demo-hp: no refusal, nothing removed (127 → 127), because every candidate was still young. The first real removals come when those copies are older than 8 days — around 11 October. demo-felhom gets its first window at its next night run. - **DooPlex keeps 8 weekly copies of ep0** (your choice). Old copies go weekly; anything ep0 deletes still never reaches the copy by itself. Today's nightly copy ran fine. - **tester-1's 3 old keys are gone.** The daily alarm for tester-1 stopped. (You got one last alarm mail this morning, sent just before I removed them.) - **A restore from the DooPlex copy works.** I restored demo-hp's whole box onto a scratch machine in about 3 minutes, read its data, then deleted it. **One trap found:** a restored box starts automatically and points at the real drives. Run beside the original, that would be two copies of the same box. I switched it off in time; the runbook now warns, and a row asks to make it safe by default. - **"Delete my old set-aside backups" works again.** The hub does it, 7 days after the household's request; the household or you can cancel in that time. Tested on tester-1 with a planted test folder. - **Your answers are recorded**, including that you keep the current Hetzner key for now (a row holds the 3 steps). - **Rows:** 6 closed, 4 opened. The list went from 330 to 328. ## Today (2026-10-03, evening): your choices A and A — built - **A box can no longer delete its off-site backups.** Both demo boxes now use a key that can only add. I tried a delete from each box: refused. Old backups all stayed (demo-felhom 11 → 13, demo-hp 91 → 100). - **No box gets the storage password any more.** The box gives the hub only its public key; the hub puts it in the storage account. I asked for the password with demo-hp's own login: refused. - **The hub keeps the passwords locked (encrypted).** A copy of the hub database no longer reveals them. - **Every day the hub checks each storage account's key file.** It found old unlocked keys from earlier boxes: 4 on demo-felhom's account and 5 on demo-hp's — now removed. tester-1's account still has 3 (its box is gone); you get a daily alarm for it until they go. - **The weekly clean-up window works, but it is switched OFF.** I opened one window on demo-hp by hand: it opened, the fake-backup check refused, and it closed in 3 seconds. The check refused because I had made a manual backup today — it is too strict. I fix that next; until then nothing deletes old backups (there is plenty of room). - **ep0's whole-box backups are copied to DooPlex every night.** First copy: 12 GB in 3 minutes, all 4 backups. DooPlex cannot read them (encrypted per household). A failed copy mails you; the test mail arrived. - **One bug, found and fixed live:** the first new box version misread its backup count as 0, and you got one false alarm mail ("demo-felhom: fell from 11 to 0"). **Ignore that mail.** Fixed 15 minutes later (0.289.1). - **Rows:** 3 closed, 6 opened, 1 opened and closed the same day. The list went from 327 to 330. ## Today (2026-10-03, later): off-site backup safety, step 1 — measured, nothing built - **The "add only" lock works.** I tested it on tester-1's storage account (your choice; no box uses it now). New backups go in. Restore works. Every delete is refused. I put the account back exactly as it was. - **But the lock alone does not protect us yet.** A broken-into box can ask the hub for the storage password. With that password it can log in and remove the lock. First the box must stop getting the password. - **A second trap:** with "add only", an attacker can add fake backups dated in the future. The normal clean-up rule then deletes all the real backups. Any clean-up must check for this. - **The hub keeps every storage password in plain form.** Anyone who reads the hub database can delete every household's off-site backups. New row. - **ep0:** Hetzner cannot snapshot the extra disk at all. One of the three old ideas does not exist. - **Rows:** 2 closed, 3 opened. The list went from 326 to 327 rows. The dated check for 6 October is done. ## Today (2026-10-03): the to-do list is in order — paperwork only, no machine touched - **Finished items left the open list.** It went from 442 rows to 326. Nothing was deleted; each moved row names where its full text is. - **Every open row now has one category and one severity** (P1 now · P2 before the first paying customer · P3 during the first customers · P4 later). **No row is P1.** 27 rows are P2. - **Two new automatic checks** refuse a finished row left in the open list, and a new row without a category or a severity. - **Your four new items are on the roadmap:** security updates for the box's own system; legal pages and business papers; "what if the household leaves Felhom"; a second login step for the dashboard. Two of them are also real findings today: **a box never receives system security updates**, and **the website has no privacy notice, terms or imprint.** - The ranked list and my reasoning: the triage recommendation in the audits folder. **Tester-2 — read only, from the hub.** The customer record exists. Tester-2's box has not registered yet. ## Today (afternoon): every app checked again - **No data-loss fault.** I checked all 58 apps with the fixed check. It made each app really save something, then looked where the data landed. **No app saves data where the backup does not copy it.** - **40 apps: proven correct. 18 apps: not proven either way.** In those 18 the check could not fill every folder (for example an upload folder stays empty because my test saves no upload). Nothing was found outside a backed-up folder. I list them and work on them later. - **papra was broken for new installs, now fixed.** It needs more memory than its limit, so a fresh install crashed in a loop. I measured it and raised the limit. No box runs papra. - **plant-it's program image is gone from Docker Hub.** It is already hidden from new installs. No box runs it. ## Removing an app now tells the truth (your choice A) - When the household removes an app, its own files (books, videos) **stay**. The dialog and the result now say so, and name the folder. Proven on the scratch box: the video was still there after the remove. ## Before the first paying customer Everything here must be done before the first customer who pays: 1. **SparkyFitness:** written permission from the author (the e-mail draft is ready). If none → hidden from new installs. 2. **Tandoor:** written permission from the authors. If none → hidden from new installs. 3. **A lawyer reviews the licence list** (Tandoor, SparkyFitness, Emby, n8n, Plex, the paid "enterprise" parts, Redis). 4. **Agent updates:** I may sign agent updates only until the first paying customer; after that you sign them. Your licence decisions are recorded: Emby, Plex and n8n stay. recipe-importer needs nothing unless we share it. ## Also today - **New version 0.288.0 and a new golden (0.288.0).** - **Rows.** 7 closed, 6 opened. The list went from 436 to 442 rows. ## What needs you 0. **The undo choice at the top of today's section** (A: keep the backup as the undo; B: build our own snapshot). If you say nothing, A stays. Earlier today's two decisions are built. The Hetzner key change stays your call. 1. **plant-it:** keep the hidden template as it is, or remove it entirely (its image no longer exists). **If you say nothing:** it stays hidden; nothing runs it. 2. **Send the SparkyFitness request, and ask the Tandoor authors** (the "Before the first paying customer" list). 3. **Phone test (2 minutes), only if you want it:** say so, and I put MeTube back on demo-hp with a family login. ## Standing steps - **Monthly security re-test: last run 2026-10-01, next due ~2026-11-01.** (You start it with the standing brief.) - **Weekly:** the golden bake (next around 9 October; today's bake was 0.288.0). - **Registry clean-up:** only when the registry disk fills; "show me" mode first, then a person decides.