Files
felhom.eu/STATUS.md
T
2026-10-05 16:25:48 +02:00

38 KiB
Raw Blame History

STATUS — what works, what's broken, what's next

Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop, off at night) was offline again; nothing was sent to it.

Updated 2026-10-05 (evening, the hub-database session): every box of ours healthy. The hub database now leaves DooPlex every night, locked, to ep0, is test-restored every Sunday, and an alarm mails you if either stops. Report: REPORT-hub-db-offsite-2026-10-05.md.

Tonight (2026-10-05, evening): the hub database off DooPlex; the cut-backup check on the scratch box

Decisions: none of mine. Yours (09 125–127): ep0 for the copy; the scratch-box restart allowed; the agent's three by-design abilities stay. Today you also chose: grow the hub's disk to 2 GiB, and restart Longhorn's disk manager.

What works now (proven live):

  • The hub makes a clean copy of its database every night at 02:00 (hub 0.136.0). The first one: 353 MB, 44 s.
  • DooPlex checks it, locks it with a key ep0 never sees, and sends it to ep0 at 02:30. First send: 7 s.
  • Every Sunday at 04:30 DooPlex takes the copy back from ep0 and checks it: it opens, it is whole, every console password in it is still locked. Done once by hand today: 4 boxes, 4 locked passwords, 0 readable.
  • The two ep0 accounts can only do their one job: the sending one cannot delete, the checking one cannot write, neither can see the households' backups. (The sending one can read back its own locked copies — that is how ep0 works.)
  • Your two saved keys work: a key rebuilt from the paper copy you saved opened the copy, and your saved lock key opened all 4 console passwords in it (a wrong key opened none).
  • The alarm: proven by Prometheus' own rule test; the real alarm mail is below.
  • A backup cut off by a restart is now said on the backup pages (scratch box): the page said so, the restore point kept the older time of the part that was not redone, and the next full backup cleared the message.

What broke, and what I did:

  • The hub's disk would not grow: Longhorn's disk manager on DooPlex was stuck. My first try (stopping the hub so the disk could grow offline) did not work and kept the hub down about 9.5 minutes. You approved restarting the disk manager: all 77 disks were back in under 2 minutes, and the hub's disk is 2 GiB now.

  • Zipline did not come back after that restart: it is set to "always the newest", so it pulled a new release that refused its database. I pinned it to the previous release; it runs and its database is updated. 7 more apps on DooPlex use "always the newest" (new row).

  • The alarm mail: I hid the "copy sent" signal on purpose; after 30 minutes the alarm fired (14:22) and the mail system sent it without error; a real send cleared it a minute later. Please check your inbox for "HubDBBackupStale" around 16:22 — I cannot read your personal mailbox.

  • The mail system (Alertmanager) cannot save its own notes since the Longhorn restart. Mail still goes out; but a silence you set would be lost at its next restart (new row).

Register: 332 → 336 rows (1 closed: the cut-backup check; 5 opened: the Longhorn fault, the "always newest" apps, a monitoring sync drift, script tests not in CI, the mail system's notes).

Needs you:

  1. Nothing urgent. If you do nothing, the copy runs every night and you get a mail only if it stops.
  2. When convenient: pin the 7 other DooPlex apps that use "always the newest" (or tell me to list them for you). If you do nothing, any restart may upgrade one of them by surprise, as it did zipline.
  3. Tester 2's one-time step is unchanged (below).

Today (2026-10-05, late afternoon): the hub's own safety; boxes left behind; the agent's admin rights

Decisions I took myself (you may reverse each — 09 decisions 119–124):

  • A box behind the approved agent for 7 days raises an alarm to you; a global version raise that cannot move a box sends you one mail naming it.
  • Scripts that post to the hub with the password must now send one extra header; a web page on another site cannot.
  • The console passwords are locked with the same key as the off-site passwords (one key to keep safe, not two).
  • The agent's rights are narrowed with exact rules and one checking helper, not one helper per command.
  • A cut-off backup is shown on the backup page until a backup runs all the way through.
  • A new agent whose root files add a file is delivered in two signed steps (the old box would refuse it in one).

What works now (proven live):

  • Form protection: a password post without the header is refused (403); a browser on another site cannot add it.
  • Console passwords locked: all 4 were sealed at the hub's start; the demo-hp one still opens its Proxmox (checked).
  • Boxes left behind: the System page lists the three per-box version floors and shows Tester 2's agent 4 releases behind. The alarm and the mail are proven by tests only (they need 7 days / a global raise).
  • The agent cannot make itself root any more: before, the real sudo let 23 of 29 attack commands through; now 0, on demo-hp and demo-felhom, and every agent feature still passes its check (67 of 67) on all three boxes.
  • Agent 0.146.1, controller 0.296.0 and hub 0.135.0 on demo-hp, demo-felhom and Tester 1; new-install image 0.296.0.

Found today:

  • The hub database is backed up — but only inside DooPlex, and only because a hand-set label says so; nothing tells anyone if that backup fails. (Your decision below.)
  • A new agent's root files could not reach any box in one step (an older box refuses files it does not know). Fixed with a two-step delivery; written down for next time.
  • A security review of my own agent change found three holes before it went to any box; fixed in a second agent release (0.146.1). Two agent releases today, against "one per repo" — the first was never sent anywhere.
  • The backup page already said "about 8 minutes", not "a few seconds". Measured today on demo-hp (9 apps): about 6 minutes. Both figures are on the page now.

Needs you:

  1. (DECIDED 2026-10-05 evening: A, done — see Tonight) Where the hub database's off-site copy goes (it holds every box's keys and your customers' settings):
    • A — my pick: ep0's backup server, encrypted on DooPlex before it leaves, with a weekly restore test and an alarm mail. Costs one small change on ep0 (a write-only account) and keeping two keys in your password manager.
    • B: a separate Hetzner Storage Box account with restic. More new parts to look after than A.
    • If you decide nothing: the database stays only on DooPlex. A fire or theft there loses every box's console password, the escrow records and the customer settings; each box would need re-pairing by hand. Steps: documentation/runbooks/RUNBOOK-hub-db-offsite-backup.md.
    • Either way, first: put the hub's lock key (OFFSITE_SECRET_KEY) in your password manager — without it a copy of the database cannot open the console passwords.
  2. (DONE 2026-10-05 evening — see Tonight) The power-cut-during-backup check (R-519) on the scratch box: the permission check refused my restarting the controller in the middle of a backup. Say "go" and the next session does it once on 9202; if not, the fix stays proven by tests only.
  3. Three things the agent can still do, by design (each written in 03 §3.1): pick which controller image its own guest runs; install the operator SSH key for the limited felhom-op user; see the box's backup key during the recovery-code ceremony. If you do nothing, they stay as they are until before the first paying customer.
  4. Tester 2's one-time step is unchanged (below).

Earlier today (2026-10-05, afternoon): a box that is not always on; the self-repair after a power cut

Decisions I took myself (you may reverse each — 09 decisions 112–118):

  • The make-up run starts 15 minutes after the box comes back; a backup due within 30 minutes is left to its normal time.
  • A laptop that sleeps through the night: a nightly job that wakes up more than an hour late is skipped (otherwise app updates would start at noon); the backups are made up instead. Tested, not measured (I may not suspend a box).
  • The banner appears when the last backup is over 26 hours old; it suggests the latest evening hour the box is usually on (5 of the last 7 days), and nothing when the box is usually on at its backup time.
  • A box that is off at the 05:00 check now raises the missed-backup alarm after 2 nights without a database backup (3 without a whole-box backup) — not after one, so a box that broke last night gives only its "offline" alarm.
  • The household hears "your server cannot be reached" at most once a week; you still hear every one.
  • The restore-test's first check is 30 minutes after the agent starts (a box on for short times now gets tested).
  • The update's repair step now also looks at dpkg's journal — the place the power cut left its mark.

What works now (proven live):

  • A missed night is made up once (your choice A): the scratch box and demo-felhom were off across their backup time; 15 minutes after they came back, the missed backups ran by themselves (seconds). The household's timeline got one line: "Kimaradt mentés pótolva: a doboz ki volt kapcsolva 10:07-kor, a mentés most elkészült." No mail.
  • The banner (your idea) appeared on the scratch box, with the "change the backup time" button; "Close" kept it closed. No screenshot: there is no browser on DooPlex; I captured the page as the box served it.
  • The power cut, again (demo-hp, your go): back by itself in 38 seconds; the next update repaired dpkg by itself and finished — nobody typed anything. No mail.
  • A restore-test 30 minutes after an agent start ran and passed on demo-felhom.
  • Controller 0.295.0, agent 0.145.0 (+ its root files) on demo-hp, demo-felhom and Tester 1; hub 0.134.0; new-install image 0.295.0 baked and approved.

Found today:

  • My slip from this morning: the Tester 1 test machine did not restart after the morning crash and stayed off for 1 h 17 min; my morning report said every box was healthy. It now starts by itself after a crash (proven by the second crash).
  • What the household may notice from a make-up run: the database backup stops an app with stored files for its copy — 1 second for opengist. The night does the same unseen; a big app may take longer, in the day (filed, small).

Needs you: nothing urgent. Tester 2's one-time step is unchanged (below). The new missed-backup alarm will be checked at tomorrow's 05:00 run; if Tester 2 is still off, you will get its first real "backup missed" mail — that is the fix working, not a new fault.

Today (2026-10-05, day): the night's fixes, the power cut by day, a box that is off at night

Decisions I took myself (you may reverse each — 09 decisions 104–108):

  • The off-site clean-up's safety line is now "the last 7 calendar days" — the same number the clean-up keeps — not "8 days old". It can never block an honest clean-up again.
  • The image clean-up now waits while ANY app install, update, restore or undo is downloading — not only installs.
  • A killed update keeps its report on the box until the hub has it; the agent looks for such reports every 5 minutes.
  • A second agent release today (0.144.1), against "one release per repo": the first fix for the lost report was proven NOT to work on demo-hp, and shipping it as it was would have been worse.

What works now (proven live):

  • The weekly off-site clean-up really deletes old backups. One clean-up each by hand: demo-felhom 16 → 14, demo-hp 145 → 127 — exactly the backups I predicted. No error mail, the hub's count check quiet, the key files clean. This also closes R-95 (the box can no longer delete its own off-site history, and clean-up now works).
  • A new box's first app install works the first time: the image clean-up met an install at minute 3 on the scratch box, waited, and BookStack installed first try. A failed install now logs its real reason.
  • The update's disk-space check counts the real download (12.8 MB for 13 packages; it counted 0 before).
  • A killed update still reports to the hub (5 minutes later, once). The debug update runs with the hub away.
  • Controller 0.294.0 on demo-hp, demo-felhom and Tester 1 (floor per customer; Tester 2 not moved). Agent 0.144.1 and its root files on all three. New-install image 0.294.0 baked and approved.

Found today (filed, not fixed):

  • After a power cut in the middle of an update, every later update fails until someone runs one command on the box (R-876, P2). The box itself comes back fine. Until the fix: runbooks/crash-guard.md has the command. I fix it next.
  • A box that is off at night (below): no catch-up, no missed-backup alarm (R-872), a "server cannot be reached" mail to the household every night (R-873), restore-tests never run (R-874).

Needs you:

  1. [DECIDED 2026-10-05 08:42 — option A, built the same afternoon; see above] A box that is off every night (Tester 2) — what does the product promise? Today such a box never gets its nightly database backups, second copy or off-site copy; the whole-box backup runs only about every 2 days; no alarm says so; the household is mailed "your server cannot be reached" every night. No design document covers it (R-871).
    • A (my pick): a missed night runs once when the box comes back. The database backups, the second copy and the off-site copy run a few minutes after the box is on again (they take seconds to minutes; whether any of them pauses an app is to be measured in the design — the whole-box backup, which does pause apps, already has its own catch-up). App and system updates still wait for a night. Costs: a design section and one controller release; the household may notice a busy disk for a few minutes after switching on.
    • B: say plainly that the box must stay on at night. The setup guide and the box's backup page say it; the missed-night alarm fires after 2 nights off. Costs: wording + one alarm; a laptop household gets an alarm it cannot fix except by changing habits.
    • If you do nothing: Tester 2 keeps having no database or off-site backup, nobody is told, and the household keeps getting the nightly "cannot be reached" mail. A restore-test alarm will fire around 2026-10-11.
  2. Tester 2's one-time step — unchanged from yesterday (below). If you wait, it keeps working; it just cannot get new root files.

Recorded, not decisions: Tester 1's Cloudflare tokens are NOT rotated (your ruling; R-870 has the steps).

Today (2026-10-04, night): root files for installed boxes, test approvals, Docker self-repair

Decisions I took myself (you may reverse each):

  • The hub cancelled only the AUTOMATIC test approvals of today (guest and host fixes). Your own "Approve Docker set" press stays in force. From now on every approval made during a test wait gets the test mark, button or not.
  • A box whose root files are behind the approved ones for 7 days sends you a mail.
  • The demo boxes' controller floor is 0.293.0; the fleet floor stays 0.292.0, so Tester 2's controller did not move.

What I did:

  • Root files for installed boxes (your ruling): a box's root-owned files (permissions list, helper scripts, crash guard) now travel as one signed package. The box checks your signature, every file and itself; on any problem it puts the old files back. Proven on both demo boxes: a wrong package refused, a package with one changed line installed and undone, a copied old job refused. New installs use the same package.
  • One catch: a box installed before tonight cannot take the FIRST package by itself; it needs one small step by hand. I did it on both demo boxes. Tester 2 needs it from you (below).
  • Tester 2 was not as old as the brief thought: it was installed at 18:06 local, after the crash guard and the new image. It already has the crash guard, live-restore and Docker 29.8.2. It lacks only tonight's Docker-update fix. I sent it the signed agent update.
  • Test approvals end with the test: the hub cancelled today's 4 test approvals at its restart (you got one mail). Tester 2 keeps what it installed; no further box installs them. The same fixes get a real approval after 24 h + a night.
  • Docker self-repair: if Docker's socket is re-created, the controller now restarts itself and the web router within about 2 minutes. Proven on the scratch box and on demo-hp, no app restarted.
  • drill-r50 is gone from the hub. On ep0 nothing was destroyed (it had no backups there); its tunnel entry left.
  • The "felhom-pbs skipped" line on Tester 2 is normal for a new box's first hour (no mail was sent).
  • New-install image 0.293.0 baked and approved.

Needs you (nothing breaks if you wait):

  • Tester 2 is offline since 20:06 local (and was off 19:13–20:05 local; it restarted in between). I sent it nothing before 20:35. My agent update for it is queued but expires at 21:20 local; when the box is back I re-send it. Worth asking the tester whether the box was switched off.
  • Tester 2's one-time step (5 minutes): connect your tunnel, ssh -p 8822 felhom-op@10.77.0.5, reveal the root password in the hub (Hosts → Tester-2 → Console access; this writes one line on Tester 2's timeline), su -, then run the three commands in documentation/runbooks/config-bundle.md ("Tester 2"). Tell me when done; I send the package. If you wait: Tester 2 keeps working; it just cannot get new root files until then.
  • Found, not fixed: the agent's permission list is wider than "minimal": a broken-into agent could become root on its own box. Worth fixing before the first paying customer.

Today (2026-10-04, ~18:30): you found a bug — the Docker update blinded the box's controller

  • What happened: the Docker update on the N100 restarted Docker itself. The apps kept running (as designed), but the controller and the web router kept a connection to the OLD Docker, so the controller could not see anything: the hub showed the N100 DOWN from 14:18 to 15:57. demo-hp had the same fault; the crash test happened to heal it.
  • Fixed (your choice): after a Docker update the box now restarts just those two (about 10 seconds, apps untouched), and the update's health check now asks "can the controller really reach Docker?". Released as host agent 0.142.1, proven twice on demo-hp, on both demo boxes now, and approved for new installs.

Today (2026-10-04, late evening): the System page, Docker updates, the crash restart

Decisions I took myself (you may reverse each):

  • The box knows a crash only as "it did not shut down cleanly" (the crash memory chip saved nothing). So a power cut also counts as a crash.
  • An "oops" (a kernel error the box survives) does not restart the box; you get a mail instead.
  • Your words win where the brief disagreed: the 3rd crash within one hour leaves the box off. Restart after 10 s, re-arm after 24 h. All settings.
  • Only the box itself decides whether a Docker update is allowed: it checks your signature with a key file only root can change (not the agent's own settings, which the agent could change).
  • The System page uses the alarm limits for its colours (red = an alarm would fire).

No decision needed from you today.

What I did:

  • The System tab in the hub: per box the Proxmox, kernel (now and next boot), Debian and Docker versions, what is waiting, held packages, "restart needed", the crash guard and the last update run — with buttons for ring, on/off, "Approve now" and "Approve Docker set". The Hosts page shows Proxmox and kernel too.
  • Docker updates: "live-restore" is on in every box (no app restarted: 24 of 24 and 5 of 5 containers kept running). Both demo boxes moved to Docker 29.8.2, every app kept running. I approved that set with the new button (a 0-night test wait, then back to 2 nights). An undo signed by your key put demo-hp back one version and forward again, apps running throughout. A copied (replayed) signed job was refused.
  • Crash restart (your 3 crashes on demo-hp): crash 1 and 2 — back by itself in under a minute; crash 3 — it stayed off until you switched it on. You got the "guard tripped" mail. I re-armed it.
  • Found: the hub learns about a crash up to 15 minutes late (nothing lost). After a crash, the household can also get an "app stopped" mail besides "restarted after a crash" — whether to calm that is a later choice for you.
  • New-install image re-made with live-restore on and the approved Docker version, and approved in the hub with host agent 0.142.0. New boxes also get the crash guard (installer 1.30.0).
  • Rows: 6 closed, 1 opened-and-closed the same day, 5 opened. The list went from 334 to 333.

Needs you later (nothing breaks if you wait):

  • Installed boxes still get new root files only by hand (the long-standing gap; there are no other boxes today).
  • Kernel updates are still not built (a hung new kernel would stay — needs a fix first).

Today (2026-10-04, evening): host fixes, the fleet view, a true tunnel status

Decisions I took myself (you may reverse each):

  • Only a box the installer itself recorded as "appliance" gets host updates (a record the agent cannot change).
  • The four alarm times: no update run for 7 days; "reboot needed" for 14 days; the demo boxes approve nothing for 7 days; a box has fixes nobody approved for 14 days. All four are settings.
  • I released the host agent and the hub a second time today. The first release said "reboot needed" wrongly; the new 14-day alarm would then have mailed you about boxes you had already rebooted.

One decision for you (a safe default if you say nothing) — the Docker engine updates (designed, not built):

  1. Turn on Docker's "live-restore" on every box. With it, a Docker engine update restarts no app (measured: 0 restarts). Without it, every app stops for about 30 seconds per engine update.
    • A (my pick): turn it on — in the new-install image and once on existing boxes. It goes on without restarting anything. Cost: it must never be turned off by a plain restart again (that stops every app and starts none).
    • B: leave it off. Every Docker update then means ~30 seconds of every app being down, at night.
    • If you say nothing: nothing changes; Docker updates stay unbuilt.

What I did:

  • The host's Debian fixes now install themselves, after the guest's, on the same night run, only on appliances, never a kernel or boot package, never a reboot. Proven: demo-felhom installed 108 host fixes and stayed healthy; all 108 came from Debian. A box made "ring 1" for a test installed exactly the one approved host version it lacked.
  • A way to put one host package back by hand is written and proven on demo-hp (and the test taught it two fixes).
  • The fleet view in the hub: one line per box with its updates, "reboot needed since", and the tunnel.
  • The tunnel status is now true: running, not running, or unknown. I blocked demo-hp's tunnel: after two reports (about 30 minutes) you got a "tunnel down" mail; when I unblocked it, "recovered". A simply stopped tunnel heals itself within 5 minutes, before the hub can even see it.
  • The night run is fast: 23–32 seconds when there is nothing to install (target was under 60).
  • The kernel test on demo-hp (your two reboots): Secure Boot works with it, but GRUB's "boot once" does not work on our boxes: the second plain reboot came up on the NEW kernel again. So a new kernel that hangs would stay. That must be solved before kernel updates. demo-hp now runs the newer kernel, healthy. A watchdog chip exists on demo-hp.
  • Found and fixed during the live test: "reboot needed" was wrong in two ways (fixed in the second release).
  • New-install image baked and vouched (0.292.0 with host agent 0.141.1). The golden waiver is gone: not needed.
  • Rows: 2 closed, 2 opened-and-closed the same day, 3 opened. The list went from 332 to 333.

Needs you later (nothing breaks if you wait):

  • A host that crashes does not restart by itself (Linux's "panic" setting is off). Changing it changes how every box behaves, so it is your call, together with kernel updates.

Today (2026-10-04, afternoon): the guest's security fixes install themselves

One decision for you (a safe default if you say nothing):

  1. Undo for a guest update that goes wrong. I tested it first, as you asked: Proxmox cannot take a snapshot of a customer box at all (the box is linked to the household's drives, and Proxmox refuses). So there is no automatic undo. Today a failed check stops, mails you, and last night's whole-box backup (minutes old) is the undo, by hand.
    • A (my pick): keep it so. Nothing new to build. A restore takes about 1–3 minutes plus losing what apps wrote since the backup.
    • B: build our own disk snapshot under Proxmox. Automatic, but a new mechanism nobody has tested, with a risk to the disk pool.
    • If you say nothing: A stays.

What I did:

  • Your two choices are built. A returning household's new box sets the old off-site copy aside on night one and starts a new one (nothing deleted). An approved version that Debian already replaced comes from Debian's dated archive.
  • Guest security fixes now install themselves. Each night, after the whole-box backup, the demo boxes install Debian's fixes. When both demo boxes run the same versions healthy for 24 hours and one night, the hub approves that set; every other box then installs exactly those versions. Docker, the host and the kernel are not touched.
    • Proven live: each demo box installed 53 fixes and stayed healthy. With a 2-minute test wait the hub approved the set; then I put back 24 hours. demo-felhom, made "ring 1" for the test, installed exactly the 3 approved versions it lacked and nothing newer. A run where I stopped an app failed its check and mailed you (you got that mail).
    • You can switch OS updates off per box in the hub. They are ON by default. The household sees one line in its timeline: "System security fixes installed".
  • The tunnel and two other built-in programs are current (cloudflared was 4 months old). A new monthly check catches them falling behind. The tunnel came back by itself on both demo boxes within 20 seconds.
  • Three bugs found and fixed during the live test, before release.
  • Two gaps found, now rows: existing boxes cannot receive the new update tool through the product (only new installs, or by hand — the demo boxes got it by hand); and the hub's "tunnel status" for every box always reads "inactive" because it checks the wrong place.
  • Rows: 4 closed (one opened and closed the same day), 5 opened. The list went from 331 to 333.

Today (2026-10-04, day): off-site closed, operating-system updates measured

Two decisions for you — each has a safe default if you say nothing:

  1. A returning household's first night. A new box for someone who had a box before makes no off-site copy, because the old copy was made with a key the new box does not have.
    • A (my pick): the new box sets the old copy aside by itself on its first night, then starts a new one. An un-claimed box already does exactly this. Nothing is deleted; you can put the old copy back. Cost: a small box change. A household that wanted to continue the old copy with its recovery code must do that before night one.
    • B: ask the household on the recovery-code evening. Cost: new screens in two languages. If they skip it, the gap stays.
    • If you say nothing: such a household has no off-site copy until someone presses "start a new off-site backup", and you get a mail each time. Nobody is blocked today (no returning household is waiting).
  2. Approved OS updates, when Debian has already replaced the version. Debian keeps only two versions of a package.
    • A (my pick): the box then fetches the exact approved version from Debian's own dated archive (snapshot.debian.org, signed by Debian). Measured: 2–3 seconds. Cost: one more outside service we rely on — but in 3 months it would have been needed zero times for the packages a box has.
    • B: the box waits for the next approval. Cost: no new service; a box can stay unpatched for one cycle.
    • If you say nothing: nothing is blocked now; the first build step can start with B and switch later.

What I did:

  • A restored box is safe by default. The automatic restore-test was already safe (measured). The agent's disaster restore now refuses to run next to a live original. Restores by hand use a new safe script: no start on boot, no link to the real drives, network off. All proven live on demo-hp. New agent 0.139.0 on both demo boxes.
  • The weekly clean-up cannot get stuck any more. You can allow one bigger clean-up for one box (hub 0.129.0). Proven on a lab copy: 98 backups, the normal limit refused, the bigger one cleaned to 13, and the next week was normal again.
  • A dated check for 12 October looks at the first clean-up that really deletes old backups.
  • OS updates, measured on the demo boxes (nothing built for customers):
    • The boxes are far behind: 188 updates per host, about 59 per guest. Every guest runs a different Docker version.
    • A Debian update of the guest took 24 seconds. No app stopped. The undo (restore the guest) took 73 seconds.
    • A Debian update of demo-hp's host took 60 seconds. The guests kept running.
    • A broken (killed) update does not fix itself. Two repair commands fix it in 5 seconds.
    • A Docker update restarts every app (about 30 seconds silent). A Docker setting ("live-restore") avoids that completely — but switching it off again stopped every app and started none. Recorded as a trap.
    • The new kernel booted fine, and the box fell back to the old kernel on the next restart. But if a new kernel hangs, the box keeps trying it. Someone must then switch it off and on. No hardware watchdog is used.
    • About 40 packages from Proxmox have ordinary names (for example the disk system ZFS and the Secure Boot loader). So "which lane" must follow where a package comes from, not its name.
    • The internet tunnel program on every box (cloudflared) is 4 months old. Nothing updates it. New row.
  • Thank you for being near the box. demo-hp restarted twice. It now runs the old kernel and starts the new one at its next restart.
  • Rows: 2 closed, 5 opened. The list went from 328 to 331.

Today (2026-10-04): off-site safety finished

  • The weekly clean-up is ON and no longer stops itself. The fake-backup check now skips young copies that a manual backup replaced the same day, instead of refusing. I ran one window on demo-hp: no refusal, nothing removed (127 → 127), because every candidate was still young. The first real removals come when those copies are older than 8 days — around 11 October. demo-felhom gets its first window at its next night run.
  • DooPlex keeps 8 weekly copies of ep0 (your choice). Old copies go weekly; anything ep0 deletes still never reaches the copy by itself. Today's nightly copy ran fine.
  • tester-1's 3 old keys are gone. The daily alarm for tester-1 stopped. (You got one last alarm mail this morning, sent just before I removed them.)
  • A restore from the DooPlex copy works. I restored demo-hp's whole box onto a scratch machine in about 3 minutes, read its data, then deleted it. One trap found: a restored box starts automatically and points at the real drives. Run beside the original, that would be two copies of the same box. I switched it off in time; the runbook now warns, and a row asks to make it safe by default.
  • "Delete my old set-aside backups" works again. The hub does it, 7 days after the household's request; the household or you can cancel in that time. Tested on tester-1 with a planted test folder.
  • Your answers are recorded, including that you keep the current Hetzner key for now (a row holds the 3 steps).
  • Rows: 6 closed, 4 opened. The list went from 330 to 328.

Today (2026-10-03, evening): your choices A and A — built

  • A box can no longer delete its off-site backups. Both demo boxes now use a key that can only add. I tried a delete from each box: refused. Old backups all stayed (demo-felhom 11 → 13, demo-hp 91 → 100).
  • No box gets the storage password any more. The box gives the hub only its public key; the hub puts it in the storage account. I asked for the password with demo-hp's own login: refused.
  • The hub keeps the passwords locked (encrypted). A copy of the hub database no longer reveals them.
  • Every day the hub checks each storage account's key file. It found old unlocked keys from earlier boxes: 4 on demo-felhom's account and 5 on demo-hp's — now removed. tester-1's account still has 3 (its box is gone); you get a daily alarm for it until they go.
  • The weekly clean-up window works, but it is switched OFF. I opened one window on demo-hp by hand: it opened, the fake-backup check refused, and it closed in 3 seconds. The check refused because I had made a manual backup today — it is too strict. I fix that next; until then nothing deletes old backups (there is plenty of room).
  • ep0's whole-box backups are copied to DooPlex every night. First copy: 12 GB in 3 minutes, all 4 backups. DooPlex cannot read them (encrypted per household). A failed copy mails you; the test mail arrived.
  • One bug, found and fixed live: the first new box version misread its backup count as 0, and you got one false alarm mail ("demo-felhom: fell from 11 to 0"). Ignore that mail. Fixed 15 minutes later (0.289.1).
  • Rows: 3 closed, 6 opened, 1 opened and closed the same day. The list went from 327 to 330.

Today (2026-10-03, later): off-site backup safety, step 1 — measured, nothing built

  • The "add only" lock works. I tested it on tester-1's storage account (your choice; no box uses it now). New backups go in. Restore works. Every delete is refused. I put the account back exactly as it was.
  • But the lock alone does not protect us yet. A broken-into box can ask the hub for the storage password. With that password it can log in and remove the lock. First the box must stop getting the password.
  • A second trap: with "add only", an attacker can add fake backups dated in the future. The normal clean-up rule then deletes all the real backups. Any clean-up must check for this.
  • The hub keeps every storage password in plain form. Anyone who reads the hub database can delete every household's off-site backups. New row.
  • ep0: Hetzner cannot snapshot the extra disk at all. One of the three old ideas does not exist.
  • Rows: 2 closed, 3 opened. The list went from 326 to 327 rows. The dated check for 6 October is done.

Today (2026-10-03): the to-do list is in order — paperwork only, no machine touched

  • Finished items left the open list. It went from 442 rows to 326. Nothing was deleted; each moved row names where its full text is.
  • Every open row now has one category and one severity (P1 now · P2 before the first paying customer · P3 during the first customers · P4 later). No row is P1. 27 rows are P2.
  • Two new automatic checks refuse a finished row left in the open list, and a new row without a category or a severity.
  • Your four new items are on the roadmap: security updates for the box's own system; legal pages and business papers; "what if the household leaves Felhom"; a second login step for the dashboard. Two of them are also real findings today: a box never receives system security updates, and the website has no privacy notice, terms or imprint.
  • The ranked list and my reasoning: the triage recommendation in the audits folder.

Tester-2 — read only, from the hub. The customer record exists. Tester-2's box has not registered yet.

Today (afternoon): every app checked again

  • No data-loss fault. I checked all 58 apps with the fixed check. It made each app really save something, then looked where the data landed. No app saves data where the backup does not copy it.
  • 40 apps: proven correct. 18 apps: not proven either way. In those 18 the check could not fill every folder (for example an upload folder stays empty because my test saves no upload). Nothing was found outside a backed-up folder. I list them and work on them later.
  • papra was broken for new installs, now fixed. It needs more memory than its limit, so a fresh install crashed in a loop. I measured it and raised the limit. No box runs papra.
  • plant-it's program image is gone from Docker Hub. It is already hidden from new installs. No box runs it.

Removing an app now tells the truth (your choice A)

  • When the household removes an app, its own files (books, videos) stay. The dialog and the result now say so, and name the folder. Proven on the scratch box: the video was still there after the remove.

Before the first paying customer

Everything here must be done before the first customer who pays:

  1. SparkyFitness: written permission from the author (the e-mail draft is ready). If none → hidden from new installs.
  2. Tandoor: written permission from the authors. If none → hidden from new installs.
  3. A lawyer reviews the licence list (Tandoor, SparkyFitness, Emby, n8n, Plex, the paid "enterprise" parts, Redis).
  4. Agent updates: I may sign agent updates only until the first paying customer; after that you sign them.

Your licence decisions are recorded: Emby, Plex and n8n stay. recipe-importer needs nothing unless we share it.

Also today

  • New version 0.288.0 and a new golden (0.288.0).
  • Rows. 7 closed, 6 opened. The list went from 436 to 442 rows.

What needs you

  1. The undo choice at the top of today's section (A: keep the backup as the undo; B: build our own snapshot). If you say nothing, A stays. Earlier today's two decisions are built. The Hetzner key change stays your call.
  2. plant-it: keep the hidden template as it is, or remove it entirely (its image no longer exists). If you say nothing: it stays hidden; nothing runs it.
  3. Send the SparkyFitness request, and ask the Tandoor authors (the "Before the first paying customer" list).
  4. Phone test (2 minutes), only if you want it: say so, and I put MeTube back on demo-hp with a family login.

Standing steps

  • Monthly security re-test: last run 2026-10-01, next due ~2026-11-01. (You start it with the standing brief.)
  • Weekly: the golden bake (next around 9 October; today's bake was 0.288.0).
  • Registry clean-up: only when the registry disk fills; "show me" mode first, then a person decides.