301fe4532a
gates / gates (push) Successful in 1m36s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
517 lines
40 KiB
Markdown
517 lines
40 KiB
Markdown
# STATUS — what works, what's broken, what's next
|
||
|
||
**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop, off at night) was offline again; nothing was
|
||
sent to it.**
|
||
|
||
**Updated 2026-10-05 (late night, burn-down round 2 — in progress): 249 open rows after your answer (was 292).**
|
||
|
||
## Your answer to the burn-down list (2026-10-05 18:23), recorded
|
||
|
||
- **43 rows closed as accepted by you**, each with its reason from the list.
|
||
- **R-124 stays open and is fixed in this session** (a recovery step that fails during a real recovery).
|
||
- **R-698 stays open as a known risk** (a backup keeps an app's image name, not the image). It is yours; nobody works on it now.
|
||
- **R-831 and R-870** (the printed tokens) stay as they are, by your earlier rulings; the rows keep their steps.
|
||
- **R-887** (the CI jobs that never ran): your screenshot shows ONE runner, online. My „old copy of the runner" guess was
|
||
wrong; the row says so.
|
||
|
||
## Tonight, later (2026-10-05, night): the list got shorter
|
||
|
||
**Decisions:** none of mine.
|
||
|
||
**What happened:**
|
||
- **Every lower-priority row was checked against today's code** (317 rows). 24 described problems a later change had
|
||
already fixed; 2 were duplicates. Those are closed, each with the change that fixed it.
|
||
- **19 small rows were fixed and closed** in four repositories — wrong comments and documents, missing tests, gates that
|
||
checked less than they claimed. No release was needed: nothing that runs on a box or on the hub changed.
|
||
- **One real fix found on the way:** the hub's build script pushed a `latest` image tag on every release. Nothing used
|
||
it; it is gone, and a test keeps it gone.
|
||
- **You got 5 „CI failed" mails tonight — my fault, fixed.** The new test gate ran one of today's test suites on the CI
|
||
machine for the first time; it uses a different `date` tool, and 9 tests failed. Fixed; CI is green again (job 1360).
|
||
- **A separate CI fault on DooPlex (new row R-887):** two CI jobs were never run and were marked failed after ~10 minutes,
|
||
with no log and no mail. (The „old runner copy" guess was wrong — see above.)
|
||
- **A new rule, so the list stops growing:** a small problem found during work (about 30 minutes) is fixed in that
|
||
session and never added to the list. Every report now states four numbers: rows before, after, opened, closed.
|
||
|
||
**The numbers:** 336 before → **292 after**; 1 opened (the CI fault below); 45 closed.
|
||
|
||
## Tonight (2026-10-05, evening): the hub database off DooPlex; the cut-backup check on the scratch box
|
||
|
||
**Decisions:** none of mine. Yours (`09` 125–127): ep0 for the copy; the scratch-box restart allowed; the agent's three
|
||
by-design abilities stay. Today you also chose: grow the hub's disk to 2 GiB, and restart Longhorn's disk manager.
|
||
|
||
**What works now (proven live):**
|
||
- **The hub makes a clean copy of its database every night at 02:00** (hub 0.136.0). The first one: 353 MB, 44 s.
|
||
- **DooPlex checks it, locks it with a key ep0 never sees, and sends it to ep0 at 02:30.** First send: 7 s.
|
||
- **Every Sunday at 04:30 DooPlex takes the copy back from ep0 and checks it**: it opens, it is whole, every console
|
||
password in it is still locked. Done once by hand today: 4 boxes, 4 locked passwords, 0 readable.
|
||
- **The two ep0 accounts can only do their one job:** the sending one cannot delete, the checking one cannot write,
|
||
neither can see the households' backups. (The sending one can read back its own locked copies — that is how ep0 works.)
|
||
- **Your two saved keys work:** a key rebuilt from the paper copy you saved opened the copy, and your saved lock key
|
||
opened all 4 console passwords in it (a wrong key opened none).
|
||
- **The alarm:** proven by Prometheus' own rule test; the real alarm mail is below.
|
||
- **A backup cut off by a restart is now said on the backup pages** (scratch box): the page said so, the restore point
|
||
kept the older time of the part that was not redone, and the next full backup cleared the message.
|
||
|
||
**What broke, and what I did:**
|
||
- **The hub's disk would not grow:** Longhorn's disk manager on DooPlex was stuck. My first try (stopping the hub so the
|
||
disk could grow offline) did not work and kept the hub **down about 9.5 minutes**. You approved restarting the disk
|
||
manager: all 77 disks were back in under 2 minutes, and the hub's disk is 2 GiB now.
|
||
- **Zipline did not come back after that restart:** it is set to "always the newest", so it pulled a new release that
|
||
refused its database. I pinned it to the previous release; it runs and its database is updated. 7 more apps on
|
||
DooPlex use "always the newest" (new row).
|
||
|
||
- **The alarm mail:** I hid the "copy sent" signal on purpose; after 30 minutes the alarm fired (14:22) and the mail
|
||
system sent it without error; a real send cleared it a minute later. **You confirmed both mails arrived:**
|
||
"[FIRING] HubDBBackupStale" 16:22 and "[RESOLVED] HubDBBackupStale" 16:27.
|
||
- **The mail system (Alertmanager) cannot save its own notes since the Longhorn restart.** Mail still goes out; but a
|
||
silence you set would be lost at its next restart (new row).
|
||
|
||
**Register:** 332 → 336 rows (1 closed: the cut-backup check; 5 opened: the Longhorn fault, the "always newest" apps, a
|
||
monitoring sync drift, script tests not in CI, the mail system's notes).
|
||
|
||
**Needs you:**
|
||
1. **Nothing urgent.** If you do nothing, the copy runs every night and you get a mail only if it stops.
|
||
2. **When convenient:** pin the 7 other DooPlex apps that use "always the newest" (or tell me to list them for you). If
|
||
you do nothing, any restart may upgrade one of them by surprise, as it did zipline.
|
||
3. **Tester 2's one-time step** is unchanged (below).
|
||
|
||
## Today (2026-10-05, late afternoon): the hub's own safety; boxes left behind; the agent's admin rights
|
||
|
||
**Decisions I took myself (you may reverse each — `09` decisions 119–124):**
|
||
- A box behind the approved agent for 7 days raises an alarm to you; a global version raise that cannot move a box
|
||
sends you one mail naming it.
|
||
- Scripts that post to the hub with the password must now send one extra header; a web page on another site cannot.
|
||
- The console passwords are locked with the same key as the off-site passwords (one key to keep safe, not two).
|
||
- The agent's rights are narrowed with exact rules and one checking helper, not one helper per command.
|
||
- A cut-off backup is shown on the backup page until a backup runs all the way through.
|
||
- A new agent whose root files add a file is delivered in two signed steps (the old box would refuse it in one).
|
||
|
||
**What works now (proven live):**
|
||
- **Form protection:** a password post without the header is refused (403); a browser on another site cannot add it.
|
||
- **Console passwords locked:** all 4 were sealed at the hub's start; the demo-hp one still opens its Proxmox (checked).
|
||
- **Boxes left behind:** the System page lists the three per-box version floors and shows Tester 2's agent 4 releases
|
||
behind. The alarm and the mail are proven by tests only (they need 7 days / a global raise).
|
||
- **The agent cannot make itself root any more:** before, the real sudo let 23 of 29 attack commands through; now 0, on
|
||
demo-hp and demo-felhom, and every agent feature still passes its check (67 of 67) on all three boxes.
|
||
- Agent 0.146.1, controller 0.296.0 and hub 0.135.0 on demo-hp, demo-felhom and Tester 1; new-install image 0.296.0.
|
||
|
||
**Found today:**
|
||
- **The hub database is backed up — but only inside DooPlex**, and only because a hand-set label says so; nothing tells
|
||
anyone if that backup fails. (Your decision below.)
|
||
- **A new agent's root files could not reach any box in one step** (an older box refuses files it does not know). Fixed
|
||
with a two-step delivery; written down for next time.
|
||
- **A security review of my own agent change found three holes** before it went to any box; fixed in a second agent
|
||
release (0.146.1). Two agent releases today, against "one per repo" — the first was never sent anywhere.
|
||
- **The backup page already said "about 8 minutes"**, not "a few seconds". Measured today on demo-hp (9 apps): about 6
|
||
minutes. Both figures are on the page now.
|
||
|
||
**Needs you:**
|
||
1. **(DECIDED 2026-10-05 evening: A, done — see Tonight)** **Where the hub database's off-site copy goes** (it holds every box's keys and your customers' settings):
|
||
- **A — my pick: ep0's backup server**, encrypted on DooPlex before it leaves, with a weekly restore test and an
|
||
alarm mail. Costs one small change on ep0 (a write-only account) and keeping two keys in your password manager.
|
||
- **B: a separate Hetzner Storage Box account** with restic. More new parts to look after than A.
|
||
- **If you decide nothing:** the database stays only on DooPlex. A fire or theft there loses every box's console
|
||
password, the escrow records and the customer settings; each box would need re-pairing by hand. Steps:
|
||
`documentation/runbooks/RUNBOOK-hub-db-offsite-backup.md`.
|
||
- **Either way, first:** put the hub's lock key (`OFFSITE_SECRET_KEY`) in your password manager — without it a copy
|
||
of the database cannot open the console passwords.
|
||
2. **(DONE 2026-10-05 evening — see Tonight)** **The power-cut-during-backup check (R-519) on the scratch box:** the permission check refused my restarting the
|
||
controller in the middle of a backup. Say "go" and the next session does it once on 9202; if not, the fix stays
|
||
proven by tests only.
|
||
3. **Three things the agent can still do, by design** (each written in `03` §3.1): pick which controller image its own
|
||
guest runs; install the operator SSH key for the limited `felhom-op` user; see the box's backup key during the
|
||
recovery-code ceremony. If you do nothing, they stay as they are until before the first paying customer.
|
||
4. **Tester 2's one-time step** is unchanged (below).
|
||
|
||
## Earlier today (2026-10-05, afternoon): a box that is not always on; the self-repair after a power cut
|
||
|
||
**Decisions I took myself (you may reverse each — `09` decisions 112–118):**
|
||
- The make-up run starts 15 minutes after the box comes back; a backup due within 30 minutes is left to its normal time.
|
||
- A laptop that sleeps through the night: a nightly job that wakes up more than an hour late is skipped (otherwise app
|
||
updates would start at noon); the backups are made up instead. Tested, not measured (I may not suspend a box).
|
||
- The banner appears when the last backup is over 26 hours old; it suggests the latest evening hour the box is usually
|
||
on (5 of the last 7 days), and nothing when the box is usually on at its backup time.
|
||
- A box that is off at the 05:00 check now raises the missed-backup alarm after 2 nights without a database backup (3
|
||
without a whole-box backup) — not after one, so a box that broke last night gives only its "offline" alarm.
|
||
- The household hears "your server cannot be reached" at most once a week; you still hear every one.
|
||
- The restore-test's first check is 30 minutes after the agent starts (a box on for short times now gets tested).
|
||
- The update's repair step now also looks at dpkg's journal — the place the power cut left its mark.
|
||
|
||
**What works now (proven live):**
|
||
- **A missed night is made up once** (your choice A): the scratch box and demo-felhom were off across their backup time;
|
||
15 minutes after they came back, the missed backups ran by themselves (seconds). The household's timeline got one line:
|
||
"Kimaradt mentés pótolva: a doboz ki volt kapcsolva 10:07-kor, a mentés most elkészült." No mail.
|
||
- **The banner** (your idea) appeared on the scratch box, with the "change the backup time" button; "Close" kept it
|
||
closed. *No screenshot: there is no browser on DooPlex; I captured the page as the box served it.*
|
||
- **The power cut, again** (demo-hp, your go): back by itself in 38 seconds; **the next update repaired dpkg by itself
|
||
and finished — nobody typed anything.** No mail.
|
||
- **A restore-test 30 minutes after an agent start** ran and passed on demo-felhom.
|
||
- Controller 0.295.0, agent 0.145.0 (+ its root files) on demo-hp, demo-felhom and Tester 1; hub 0.134.0; new-install
|
||
image 0.295.0 baked and approved.
|
||
|
||
**Found today:**
|
||
- **My slip from this morning:** the Tester 1 test machine did not restart after the morning crash and stayed off for
|
||
1 h 17 min; my morning report said every box was healthy. It now starts by itself after a crash (proven by the second
|
||
crash).
|
||
- **What the household may notice from a make-up run:** the database backup stops an app with stored files for its copy —
|
||
1 second for opengist. The night does the same unseen; a big app may take longer, in the day (filed, small).
|
||
|
||
**Needs you:** nothing urgent. **Tester 2's one-time step** is unchanged (below). The new missed-backup alarm will be
|
||
checked at tomorrow's 05:00 run; if Tester 2 is still off, you will get its first real "backup missed" mail — that is
|
||
the fix working, not a new fault.
|
||
|
||
## Today (2026-10-05, day): the night's fixes, the power cut by day, a box that is off at night
|
||
|
||
**Decisions I took myself (you may reverse each — `09` decisions 104–108):**
|
||
- The off-site clean-up's safety line is now "the last 7 calendar days" — the same number the clean-up keeps — not
|
||
"8 days old". It can never block an honest clean-up again.
|
||
- The image clean-up now waits while ANY app install, update, restore or undo is downloading — not only installs.
|
||
- A killed update keeps its report on the box until the hub has it; the agent looks for such reports every 5 minutes.
|
||
- **A second agent release today (0.144.1)**, against "one release per repo": the first fix for the lost report was
|
||
proven NOT to work on demo-hp, and shipping it as it was would have been worse.
|
||
|
||
**What works now (proven live):**
|
||
- **The weekly off-site clean-up really deletes old backups.** One clean-up each by hand: demo-felhom 16 → 14, demo-hp
|
||
145 → 127 — exactly the backups I predicted. No error mail, the hub's count check quiet, the key files clean.
|
||
This also closes R-95 (the box can no longer delete its own off-site history, and clean-up now works).
|
||
- **A new box's first app install works the first time:** the image clean-up met an install at minute 3 on the scratch
|
||
box, waited, and BookStack installed first try. A failed install now logs its real reason.
|
||
- **The update's disk-space check counts the real download** (12.8 MB for 13 packages; it counted 0 before).
|
||
- **A killed update still reports to the hub** (5 minutes later, once). The debug update runs with the hub away.
|
||
- Controller 0.294.0 on demo-hp, demo-felhom and Tester 1 (floor per customer; Tester 2 not moved). Agent 0.144.1
|
||
and its root files on all three. New-install image 0.294.0 baked and approved.
|
||
|
||
**Found today (filed, not fixed):**
|
||
- **After a power cut in the middle of an update, every later update fails until someone runs one command on the box**
|
||
(R-876, P2). The box itself comes back fine. Until the fix: `runbooks/crash-guard.md` has the command. I fix it next.
|
||
- A box that is off at night (below): no catch-up, no missed-backup alarm (R-872), a "server cannot be reached" mail to
|
||
the household every night (R-873), restore-tests never run (R-874).
|
||
|
||
**Needs you:**
|
||
|
||
1. **[DECIDED 2026-10-05 08:42 — option A, built the same afternoon; see above]** **A box that is off every night (Tester 2) — what does the product promise?** Today such a box never gets its
|
||
nightly database backups, second copy or off-site copy; the whole-box backup runs only about every 2 days; no
|
||
alarm says so; the household is mailed "your server cannot be reached" every night. No design document covers it
|
||
(R-871).
|
||
- **A (my pick): a missed night runs once when the box comes back.** The database backups, the second copy and
|
||
the off-site copy run a few minutes after the box is on again (they take seconds to minutes; whether
|
||
any of them pauses an app is to be measured in the design — the whole-box backup, which does pause apps, already
|
||
has its own catch-up). App and system updates still
|
||
wait for a night. Costs: a design section and one controller release; the household may notice a busy disk for a
|
||
few minutes after switching on.
|
||
- **B: say plainly that the box must stay on at night.** The setup guide and the box's backup page say it; the
|
||
missed-night alarm fires after 2 nights off. Costs: wording + one alarm; a laptop household gets an alarm it
|
||
cannot fix except by changing habits.
|
||
- **If you do nothing:** Tester 2 keeps having no database or off-site backup, nobody is told, and the household
|
||
keeps getting the nightly "cannot be reached" mail. A restore-test alarm will fire around 2026-10-11.
|
||
2. **Tester 2's one-time step** — unchanged from yesterday (below). If you wait, it keeps working; it just cannot get
|
||
new root files.
|
||
|
||
**Recorded, not decisions:** Tester 1's Cloudflare tokens are NOT rotated (your ruling; R-870 has the steps).
|
||
|
||
## Today (2026-10-04, night): root files for installed boxes, test approvals, Docker self-repair
|
||
|
||
**Decisions I took myself (you may reverse each):**
|
||
- The hub cancelled only the AUTOMATIC test approvals of today (guest and host fixes). Your own "Approve Docker set"
|
||
press stays in force. From now on every approval made during a test wait gets the test mark, button or not.
|
||
- A box whose root files are behind the approved ones for 7 days sends you a mail.
|
||
- The demo boxes' controller floor is 0.293.0; the fleet floor stays 0.292.0, so Tester 2's controller did not move.
|
||
|
||
**What I did:**
|
||
- **Root files for installed boxes (your ruling):** a box's root-owned files (permissions list, helper scripts, crash
|
||
guard) now travel as one signed package. The box checks your signature, every file and itself; on any problem it
|
||
puts the old files back. Proven on both demo boxes: a wrong package refused, a package with one changed line
|
||
installed and undone, a copied old job refused. New installs use the same package.
|
||
- **One catch:** a box installed before tonight cannot take the FIRST package by itself; it needs one small step by
|
||
hand. I did it on both demo boxes. **Tester 2 needs it from you** (below).
|
||
- **Tester 2 was not as old as the brief thought:** it was installed at 18:06 local, after the crash guard and the new
|
||
image. It already has the crash guard, live-restore and Docker 29.8.2. It lacks only tonight's Docker-update fix. I
|
||
sent it the signed agent update.
|
||
- **Test approvals end with the test:** the hub cancelled today's 4 test approvals at its restart (you got one mail).
|
||
Tester 2 keeps what it installed; no further box installs them. The same fixes get a real approval after 24 h + a night.
|
||
- **Docker self-repair:** if Docker's socket is re-created, the controller now restarts itself and the web router within
|
||
about 2 minutes. Proven on the scratch box and on demo-hp, no app restarted.
|
||
- **drill-r50 is gone** from the hub. On ep0 nothing was destroyed (it had no backups there); its tunnel entry left.
|
||
- The "felhom-pbs skipped" line on Tester 2 is normal for a new box's first hour (no mail was sent).
|
||
- New-install image 0.293.0 baked and approved.
|
||
|
||
**Needs you (nothing breaks if you wait):**
|
||
- **Tester 2 is offline** since 20:06 local (and was off 19:13–20:05 local; it restarted in between). I sent it nothing
|
||
before 20:35. My agent update for it is queued but expires at 21:20 local; when the box is back I re-send it.
|
||
Worth asking the tester whether the box was switched off.
|
||
- **Tester 2's one-time step** (5 minutes): connect your tunnel, `ssh -p 8822 felhom-op@10.77.0.5`, reveal the root
|
||
password in the hub (Hosts → Tester-2 → Console access; this writes one line on Tester 2's timeline), `su -`, then
|
||
run the three commands in `documentation/runbooks/config-bundle.md` ("Tester 2"). Tell me when done; I send the package.
|
||
If you wait: Tester 2 keeps working; it just cannot get new root files until then.
|
||
- **Found, not fixed:** the agent's permission list is wider than "minimal": a broken-into agent could become root on
|
||
its own box. Worth fixing before the first paying customer.
|
||
|
||
## Today (2026-10-04, ~18:30): you found a bug — the Docker update blinded the box's controller
|
||
|
||
- **What happened:** the Docker update on the N100 restarted Docker itself. The apps kept running (as designed), but
|
||
the controller and the web router kept a connection to the OLD Docker, so the controller could not see anything:
|
||
the hub showed the N100 DOWN from 14:18 to 15:57. demo-hp had the same fault; the crash test happened to heal it.
|
||
- **Fixed (your choice):** after a Docker update the box now restarts just those two (about 10 seconds, apps untouched),
|
||
and the update's health check now asks "can the controller really reach Docker?". Released as host agent 0.142.1,
|
||
proven twice on demo-hp, on both demo boxes now, and approved for new installs.
|
||
|
||
## Today (2026-10-04, late evening): the System page, Docker updates, the crash restart
|
||
|
||
**Decisions I took myself (you may reverse each):**
|
||
- The box knows a crash only as "it did not shut down cleanly" (the crash memory chip saved nothing). So a power cut
|
||
also counts as a crash.
|
||
- An "oops" (a kernel error the box survives) does not restart the box; you get a mail instead.
|
||
- Your words win where the brief disagreed: the **3rd** crash within one hour leaves the box off. Restart after 10 s,
|
||
re-arm after 24 h. All settings.
|
||
- Only the box itself decides whether a Docker update is allowed: it checks your signature with a key file only root
|
||
can change (not the agent's own settings, which the agent could change).
|
||
- The System page uses the alarm limits for its colours (red = an alarm would fire).
|
||
|
||
**No decision needed from you today.**
|
||
|
||
**What I did:**
|
||
- **The System tab** in the hub: per box the Proxmox, kernel (now and next boot), Debian and Docker versions, what is
|
||
waiting, held packages, "restart needed", the crash guard and the last update run — with buttons for ring, on/off,
|
||
"Approve now" and "Approve Docker set". The Hosts page shows Proxmox and kernel too.
|
||
- **Docker updates:** "live-restore" is on in every box (no app restarted: 24 of 24 and 5 of 5 containers kept running).
|
||
Both demo boxes moved to Docker 29.8.2, every app kept running. I approved that set with the new button (a 0-night
|
||
test wait, then back to 2 nights). An undo signed by your key put demo-hp back one version and forward again, apps
|
||
running throughout. A copied (replayed) signed job was refused.
|
||
- **Crash restart (your 3 crashes on demo-hp):** crash 1 and 2 — back by itself in under a minute; crash 3 — it stayed
|
||
off until you switched it on. You got the "guard tripped" mail. I re-armed it.
|
||
- **Found:** the hub learns about a crash up to 15 minutes late (nothing lost). After a crash, the household can also get
|
||
an "app stopped" mail besides "restarted after a crash" — whether to calm that is a later choice for you.
|
||
- **New-install image re-made** with live-restore on and the approved Docker version, and approved in the hub with host
|
||
agent 0.142.0. New boxes also get the crash guard (installer 1.30.0).
|
||
- **Rows:** 6 closed, 1 opened-and-closed the same day, 5 opened. The list went from 334 to 333.
|
||
|
||
**Needs you later (nothing breaks if you wait):**
|
||
- Installed boxes still get new root files only by hand (the long-standing gap; there are no other boxes today).
|
||
- Kernel updates are still not built (a hung new kernel would stay — needs a fix first).
|
||
|
||
## Today (2026-10-04, evening): host fixes, the fleet view, a true tunnel status
|
||
|
||
**Decisions I took myself (you may reverse each):**
|
||
- Only a box the installer itself recorded as "appliance" gets host updates (a record the agent cannot change).
|
||
- The four alarm times: no update run for 7 days; "reboot needed" for 14 days; the demo boxes approve nothing for 7
|
||
days; a box has fixes nobody approved for 14 days. All four are settings.
|
||
- I released the host agent and the hub a second time today. The first release said "reboot needed" wrongly; the new
|
||
14-day alarm would then have mailed you about boxes you had already rebooted.
|
||
|
||
**One decision for you (a safe default if you say nothing) — the Docker engine updates (designed, not built):**
|
||
1. **Turn on Docker's "live-restore" on every box.** With it, a Docker engine update restarts no app (measured: 0
|
||
restarts). Without it, every app stops for about 30 seconds per engine update.
|
||
- **A (my pick):** turn it on — in the new-install image and once on existing boxes. It goes on without restarting
|
||
anything. Cost: it must never be turned off by a plain restart again (that stops every app and starts none).
|
||
- **B:** leave it off. Every Docker update then means ~30 seconds of every app being down, at night.
|
||
- **If you say nothing:** nothing changes; Docker updates stay unbuilt.
|
||
|
||
**What I did:**
|
||
- **The host's Debian fixes now install themselves**, after the guest's, on the same night run, only on appliances,
|
||
never a kernel or boot package, never a reboot. Proven: demo-felhom installed 108 host fixes and stayed healthy; all
|
||
108 came from Debian. A box made "ring 1" for a test installed exactly the one approved host version it lacked.
|
||
- **A way to put one host package back by hand** is written and proven on demo-hp (and the test taught it two fixes).
|
||
- **The fleet view** in the hub: one line per box with its updates, "reboot needed since", and the tunnel.
|
||
- **The tunnel status is now true:** running, not running, or unknown. I blocked demo-hp's tunnel: after two reports
|
||
(about 30 minutes) you got a "tunnel down" mail; when I unblocked it, "recovered". A simply stopped tunnel heals
|
||
itself within 5 minutes, before the hub can even see it.
|
||
- **The night run is fast:** 23–32 seconds when there is nothing to install (target was under 60).
|
||
- **The kernel test on demo-hp (your two reboots):** Secure Boot works with it, but GRUB's "boot once" does not work
|
||
on our boxes: the second plain reboot came up on the NEW kernel again. So a new kernel that hangs would stay. That
|
||
must be solved before kernel updates. demo-hp now runs the newer kernel, healthy. A watchdog chip exists on demo-hp.
|
||
- **Found and fixed during the live test:** "reboot needed" was wrong in two ways (fixed in the second release).
|
||
- **New-install image baked and vouched** (0.292.0 with host agent 0.141.1). The golden waiver is gone: not needed.
|
||
- **Rows:** 2 closed, 2 opened-and-closed the same day, 3 opened. The list went from 332 to 333.
|
||
|
||
**Needs you later (nothing breaks if you wait):**
|
||
- **A host that crashes does not restart by itself** (Linux's "panic" setting is off). Changing it changes how every
|
||
box behaves, so it is your call, together with kernel updates.
|
||
|
||
## Today (2026-10-04, afternoon): the guest's security fixes install themselves
|
||
|
||
**One decision for you (a safe default if you say nothing):**
|
||
|
||
1. **Undo for a guest update that goes wrong.** I tested it first, as you asked: Proxmox cannot take a snapshot of a
|
||
customer box at all (the box is linked to the household's drives, and Proxmox refuses). So there is no automatic undo.
|
||
Today a failed check stops, mails you, and last night's whole-box backup (minutes old) is the undo, by hand.
|
||
- **A (my pick):** keep it so. Nothing new to build. A restore takes about 1–3 minutes plus losing what apps wrote
|
||
since the backup.
|
||
- **B:** build our own disk snapshot under Proxmox. Automatic, but a new mechanism nobody has tested, with a risk to
|
||
the disk pool.
|
||
- **If you say nothing:** A stays.
|
||
|
||
**What I did:**
|
||
- **Your two choices are built.** A returning household's new box sets the old off-site copy aside on night one and
|
||
starts a new one (nothing deleted). An approved version that Debian already replaced comes from Debian's dated archive.
|
||
- **Guest security fixes now install themselves.** Each night, after the whole-box backup, the demo boxes install
|
||
Debian's fixes. When both demo boxes run the same versions healthy for 24 hours and one night, the hub approves that
|
||
set; every other box then installs exactly those versions. Docker, the host and the kernel are not touched.
|
||
- Proven live: each demo box installed 53 fixes and stayed healthy. With a 2-minute test wait the hub approved the
|
||
set; then I put back 24 hours. demo-felhom, made "ring 1" for the test, installed exactly the 3 approved versions it
|
||
lacked and nothing newer. A run where I stopped an app failed its check and mailed you (you got that mail).
|
||
- You can switch OS updates off per box in the hub. They are ON by default. The household sees one line in its
|
||
timeline: "System security fixes installed".
|
||
- **The tunnel and two other built-in programs are current** (cloudflared was 4 months old). A new monthly check
|
||
catches them falling behind. The tunnel came back by itself on both demo boxes within 20 seconds.
|
||
- **Three bugs found and fixed during the live test**, before release.
|
||
- **Two gaps found, now rows:** existing boxes cannot receive the new update tool through the product (only new
|
||
installs, or by hand — the demo boxes got it by hand); and the hub's "tunnel status" for every box always reads
|
||
"inactive" because it checks the wrong place.
|
||
- **Rows:** 4 closed (one opened and closed the same day), 5 opened. The list went from 331 to 333.
|
||
|
||
## Today (2026-10-04, day): off-site closed, operating-system updates measured
|
||
|
||
**Two decisions for you — each has a safe default if you say nothing:**
|
||
|
||
1. **A returning household's first night.** A new box for someone who had a box before makes no off-site copy, because
|
||
the old copy was made with a key the new box does not have.
|
||
- **A (my pick):** the new box sets the old copy aside by itself on its first night, then starts a new one. An
|
||
un-claimed box already does exactly this. Nothing is deleted; you can put the old copy back. Cost: a small box
|
||
change. A household that wanted to continue the old copy with its recovery code must do that before night one.
|
||
- **B:** ask the household on the recovery-code evening. Cost: new screens in two languages. If they skip it, the gap stays.
|
||
- **If you say nothing:** such a household has no off-site copy until someone presses "start a new off-site backup",
|
||
and you get a mail each time. Nobody is blocked today (no returning household is waiting).
|
||
2. **Approved OS updates, when Debian has already replaced the version.** Debian keeps only two versions of a package.
|
||
- **A (my pick):** the box then fetches the exact approved version from Debian's own dated archive
|
||
(snapshot.debian.org, signed by Debian). Measured: 2–3 seconds. Cost: one more outside service we rely on — but in
|
||
3 months it would have been needed **zero** times for the packages a box has.
|
||
- **B:** the box waits for the next approval. Cost: no new service; a box can stay unpatched for one cycle.
|
||
- **If you say nothing:** nothing is blocked now; the first build step can start with B and switch later.
|
||
|
||
**What I did:**
|
||
- **A restored box is safe by default.** The automatic restore-test was already safe (measured). The agent's disaster
|
||
restore now refuses to run next to a live original. Restores by hand use a new safe script: no start on boot, no link
|
||
to the real drives, network off. All proven live on demo-hp. New agent 0.139.0 on both demo boxes.
|
||
- **The weekly clean-up cannot get stuck any more.** You can allow one bigger clean-up for one box (hub 0.129.0). Proven on
|
||
a lab copy: 98 backups, the normal limit refused, the bigger one cleaned to 13, and the next week was normal again.
|
||
- **A dated check for 12 October** looks at the first clean-up that really deletes old backups.
|
||
- **OS updates, measured on the demo boxes (nothing built for customers):**
|
||
- The boxes are far behind: 188 updates per host, about 59 per guest. Every guest runs a different Docker version.
|
||
- A Debian update of the guest took 24 seconds. No app stopped. The undo (restore the guest) took 73 seconds.
|
||
- A Debian update of demo-hp's host took 60 seconds. The guests kept running.
|
||
- A broken (killed) update does not fix itself. Two repair commands fix it in 5 seconds.
|
||
- A Docker update restarts every app (about 30 seconds silent). A Docker setting ("live-restore") avoids that
|
||
completely — but switching it off again stopped every app and started none. Recorded as a trap.
|
||
- The new kernel booted fine, and the box fell back to the old kernel on the next restart. **But** if a new kernel
|
||
hangs, the box keeps trying it. Someone must then switch it off and on. No hardware watchdog is used.
|
||
- About 40 packages from Proxmox have ordinary names (for example the disk system ZFS and the Secure Boot loader). So
|
||
"which lane" must follow where a package comes from, not its name.
|
||
- The internet tunnel program on every box (cloudflared) is 4 months old. Nothing updates it. New row.
|
||
- **Thank you for being near the box.** demo-hp restarted twice. It now runs the old kernel and starts the new one at
|
||
its next restart.
|
||
- **Rows:** 2 closed, 5 opened. The list went from 328 to 331.
|
||
|
||
## Today (2026-10-04): off-site safety finished
|
||
|
||
- **The weekly clean-up is ON and no longer stops itself.** The fake-backup check now skips young copies that a manual
|
||
backup replaced the same day, instead of refusing. I ran one window on demo-hp: no refusal, nothing removed
|
||
(127 → 127), because every candidate was still young. The first real removals come when those copies are older than
|
||
8 days — around 11 October. demo-felhom gets its first window at its next night run.
|
||
- **DooPlex keeps 8 weekly copies of ep0** (your choice). Old copies go weekly; anything ep0 deletes still never reaches
|
||
the copy by itself. Today's nightly copy ran fine.
|
||
- **tester-1's 3 old keys are gone.** The daily alarm for tester-1 stopped. (You got one last alarm mail this morning,
|
||
sent just before I removed them.)
|
||
- **A restore from the DooPlex copy works.** I restored demo-hp's whole box onto a scratch machine in about 3 minutes,
|
||
read its data, then deleted it. **One trap found:** a restored box starts automatically and points at the real
|
||
drives. Run beside the original, that would be two copies of the same box. I switched it off in time; the runbook
|
||
now warns, and a row asks to make it safe by default.
|
||
- **"Delete my old set-aside backups" works again.** The hub does it, 7 days after the household's request; the
|
||
household or you can cancel in that time. Tested on tester-1 with a planted test folder.
|
||
- **Your answers are recorded**, including that you keep the current Hetzner key for now (a row holds the 3 steps).
|
||
- **Rows:** 6 closed, 4 opened. The list went from 330 to 328.
|
||
|
||
## Today (2026-10-03, evening): your choices A and A — built
|
||
|
||
- **A box can no longer delete its off-site backups.** Both demo boxes now use a key that can only add. I tried a
|
||
delete from each box: refused. Old backups all stayed (demo-felhom 11 → 13, demo-hp 91 → 100).
|
||
- **No box gets the storage password any more.** The box gives the hub only its public key; the hub puts it in the
|
||
storage account. I asked for the password with demo-hp's own login: refused.
|
||
- **The hub keeps the passwords locked (encrypted).** A copy of the hub database no longer reveals them.
|
||
- **Every day the hub checks each storage account's key file.** It found old unlocked keys from earlier boxes:
|
||
4 on demo-felhom's account and 5 on demo-hp's — now removed. tester-1's account still has 3 (its box is gone); you
|
||
get a daily alarm for it until they go.
|
||
- **The weekly clean-up window works, but it is switched OFF.** I opened one window on demo-hp by hand: it opened,
|
||
the fake-backup check refused, and it closed in 3 seconds. The check refused because I had made a manual backup
|
||
today — it is too strict. I fix that next; until then nothing deletes old backups (there is plenty of room).
|
||
- **ep0's whole-box backups are copied to DooPlex every night.** First copy: 12 GB in 3 minutes, all 4 backups.
|
||
DooPlex cannot read them (encrypted per household). A failed copy mails you; the test mail arrived.
|
||
- **One bug, found and fixed live:** the first new box version misread its backup count as 0, and you got one
|
||
false alarm mail ("demo-felhom: fell from 11 to 0"). **Ignore that mail.** Fixed 15 minutes later (0.289.1).
|
||
- **Rows:** 3 closed, 6 opened, 1 opened and closed the same day. The list went from 327 to 330.
|
||
|
||
## Today (2026-10-03, later): off-site backup safety, step 1 — measured, nothing built
|
||
|
||
- **The "add only" lock works.** I tested it on tester-1's storage account (your choice; no box uses it now).
|
||
New backups go in. Restore works. Every delete is refused. I put the account back exactly as it was.
|
||
- **But the lock alone does not protect us yet.** A broken-into box can ask the hub for the storage password.
|
||
With that password it can log in and remove the lock. First the box must stop getting the password.
|
||
- **A second trap:** with "add only", an attacker can add fake backups dated in the future. The normal
|
||
clean-up rule then deletes all the real backups. Any clean-up must check for this.
|
||
- **The hub keeps every storage password in plain form.** Anyone who reads the hub database can delete
|
||
every household's off-site backups. New row.
|
||
- **ep0:** Hetzner cannot snapshot the extra disk at all. One of the three old ideas does not exist.
|
||
- **Rows:** 2 closed, 3 opened. The list went from 326 to 327 rows. The dated check for 6 October is done.
|
||
|
||
## Today (2026-10-03): the to-do list is in order — paperwork only, no machine touched
|
||
|
||
- **Finished items left the open list.** It went from 442 rows to 326. Nothing was deleted; each moved row names
|
||
where its full text is.
|
||
- **Every open row now has one category and one severity** (P1 now · P2 before the first paying customer · P3
|
||
during the first customers · P4 later). **No row is P1.** 27 rows are P2.
|
||
- **Two new automatic checks** refuse a finished row left in the open list, and a new row without a category or a
|
||
severity.
|
||
- **Your four new items are on the roadmap:** security updates for the box's own system; legal pages and business
|
||
papers; "what if the household leaves Felhom"; a second login step for the dashboard. Two of them are also real
|
||
findings today: **a box never receives system security updates**, and **the website has no privacy notice,
|
||
terms or imprint.**
|
||
- The ranked list and my reasoning: the triage recommendation in the audits folder.
|
||
|
||
**Tester-2 — read only, from the hub.** The customer record exists. Tester-2's box has not registered yet.
|
||
|
||
## Today (afternoon): every app checked again
|
||
|
||
- **No data-loss fault.** I checked all 58 apps with the fixed check. It made each app really save something, then
|
||
looked where the data landed. **No app saves data where the backup does not copy it.**
|
||
- **40 apps: proven correct. 18 apps: not proven either way.** In those 18 the check could not fill every folder (for
|
||
example an upload folder stays empty because my test saves no upload). Nothing was found outside a backed-up folder.
|
||
I list them and work on them later.
|
||
- **papra was broken for new installs, now fixed.** It needs more memory than its limit, so a fresh install crashed in a
|
||
loop. I measured it and raised the limit. No box runs papra.
|
||
- **plant-it's program image is gone from Docker Hub.** It is already hidden from new installs. No box runs it.
|
||
|
||
## Removing an app now tells the truth (your choice A)
|
||
|
||
- When the household removes an app, its own files (books, videos) **stay**. The dialog and the result now say so, and
|
||
name the folder. Proven on the scratch box: the video was still there after the remove.
|
||
|
||
## Before the first paying customer
|
||
|
||
Everything here must be done before the first customer who pays:
|
||
|
||
1. **SparkyFitness:** written permission from the author (the e-mail draft is ready). If none → hidden from new installs.
|
||
2. **Tandoor:** written permission from the authors. If none → hidden from new installs.
|
||
3. **A lawyer reviews the licence list** (Tandoor, SparkyFitness, Emby, n8n, Plex, the paid "enterprise" parts, Redis).
|
||
4. **Agent updates:** I may sign agent updates only until the first paying customer; after that you sign them.
|
||
|
||
Your licence decisions are recorded: Emby, Plex and n8n stay. recipe-importer needs nothing unless we share it.
|
||
|
||
## Also today
|
||
|
||
- **New version 0.288.0 and a new golden (0.288.0).**
|
||
- **Rows.** 7 closed, 6 opened. The list went from 436 to 442 rows.
|
||
|
||
## What needs you
|
||
|
||
0. **The undo choice at the top of today's section** (A: keep the backup as the undo; B: build our own snapshot).
|
||
If you say nothing, A stays. Earlier today's two decisions are built. The Hetzner key change stays your call.
|
||
1. **plant-it:** keep the hidden template as it is, or remove it entirely (its image no longer exists). **If you say
|
||
nothing:** it stays hidden; nothing runs it.
|
||
2. **Send the SparkyFitness request, and ask the Tandoor authors** (the "Before the first paying customer" list).
|
||
3. **Phone test (2 minutes), only if you want it:** say so, and I put MeTube back on demo-hp with a family login.
|
||
|
||
## Standing steps
|
||
|
||
- **Monthly security re-test: last run 2026-10-01, next due ~2026-11-01.** (You start it with the standing brief.)
|
||
- **Weekly:** the golden bake (next around 9 October; today's bake was 0.288.0).
|
||
- **Registry clean-up:** only when the registry disk fills; "show me" mode first, then a person decides.
|