Files
felhom.eu/STATUS.md
T

195 lines
15 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# STATUS — what works, what's broken, what's next
**Ready for the first real tester (Tester-2): yes. You confirmed the tunnel route and the connect mails (2026-09-30).**
**Updated 2026-10-04 (afternoon): guest system updates are automatic on the demo boxes. Both demo boxes run
controller 0.291.0 and host agent 0.140.0. Hub 0.130.0. New installs get golden 0.291.0.**
## Today (2026-10-04, afternoon): the guest's security fixes install themselves
**One decision for you (a safe default if you say nothing):**
1. **Undo for a guest update that goes wrong.** I tested it first, as you asked: Proxmox cannot take a snapshot of a
customer box at all (the box is linked to the household's drives, and Proxmox refuses). So there is no automatic undo.
Today a failed check stops, mails you, and last night's whole-box backup (minutes old) is the undo, by hand.
- **A (my pick):** keep it so. Nothing new to build. A restore takes about 1–3 minutes plus losing what apps wrote
since the backup.
- **B:** build our own disk snapshot under Proxmox. Automatic, but a new mechanism nobody has tested, with a risk to
the disk pool.
- **If you say nothing:** A stays.
**What I did:**
- **Your two choices are built.** A returning household's new box sets the old off-site copy aside on night one and
starts a new one (nothing deleted). An approved version that Debian already replaced comes from Debian's dated archive.
- **Guest security fixes now install themselves.** Each night, after the whole-box backup, the demo boxes install
Debian's fixes. When both demo boxes run the same versions healthy for 24 hours and one night, the hub approves that
set; every other box then installs exactly those versions. Docker, the host and the kernel are not touched.
- Proven live: each demo box installed 53 fixes and stayed healthy. With a 2-minute test wait the hub approved the
set; then I put back 24 hours. demo-felhom, made "ring 1" for the test, installed exactly the 3 approved versions it
lacked and nothing newer. A run where I stopped an app failed its check and mailed you (you got that mail).
- You can switch OS updates off per box in the hub. They are ON by default. The household sees one line in its
timeline: "System security fixes installed".
- **The tunnel and two other built-in programs are current** (cloudflared was 4 months old). A new monthly check
catches them falling behind. The tunnel came back by itself on both demo boxes within 20 seconds.
- **Three bugs found and fixed during the live test**, before release.
- **Two gaps found, now rows:** existing boxes cannot receive the new update tool through the product (only new
installs, or by hand — the demo boxes got it by hand); and the hub's "tunnel status" for every box always reads
"inactive" because it checks the wrong place.
- **Rows:** 4 closed (one opened and closed the same day), 5 opened. The list went from 331 to 333.
## Today (2026-10-04, day): off-site closed, operating-system updates measured
**Two decisions for you — each has a safe default if you say nothing:**
1. **A returning household's first night.** A new box for someone who had a box before makes no off-site copy, because
the old copy was made with a key the new box does not have.
- **A (my pick):** the new box sets the old copy aside by itself on its first night, then starts a new one. An
un-claimed box already does exactly this. Nothing is deleted; you can put the old copy back. Cost: a small box
change. A household that wanted to continue the old copy with its recovery code must do that before night one.
- **B:** ask the household on the recovery-code evening. Cost: new screens in two languages. If they skip it, the gap stays.
- **If you say nothing:** such a household has no off-site copy until someone presses "start a new off-site backup",
and you get a mail each time. Nobody is blocked today (no returning household is waiting).
2. **Approved OS updates, when Debian has already replaced the version.** Debian keeps only two versions of a package.
- **A (my pick):** the box then fetches the exact approved version from Debian's own dated archive
(snapshot.debian.org, signed by Debian). Measured: 2–3 seconds. Cost: one more outside service we rely on — but in
3 months it would have been needed **zero** times for the packages a box has.
- **B:** the box waits for the next approval. Cost: no new service; a box can stay unpatched for one cycle.
- **If you say nothing:** nothing is blocked now; the first build step can start with B and switch later.
**What I did:**
- **A restored box is safe by default.** The automatic restore-test was already safe (measured). The agent's disaster
restore now refuses to run next to a live original. Restores by hand use a new safe script: no start on boot, no link
to the real drives, network off. All proven live on demo-hp. New agent 0.139.0 on both demo boxes.
- **The weekly clean-up cannot get stuck any more.** You can allow one bigger clean-up for one box (hub 0.129.0). Proven on
a lab copy: 98 backups, the normal limit refused, the bigger one cleaned to 13, and the next week was normal again.
- **A dated check for 12 October** looks at the first clean-up that really deletes old backups.
- **OS updates, measured on the demo boxes (nothing built for customers):**
- The boxes are far behind: 188 updates per host, about 59 per guest. Every guest runs a different Docker version.
- A Debian update of the guest took 24 seconds. No app stopped. The undo (restore the guest) took 73 seconds.
- A Debian update of demo-hp's host took 60 seconds. The guests kept running.
- A broken (killed) update does not fix itself. Two repair commands fix it in 5 seconds.
- A Docker update restarts every app (about 30 seconds silent). A Docker setting ("live-restore") avoids that
completely — but switching it off again stopped every app and started none. Recorded as a trap.
- The new kernel booted fine, and the box fell back to the old kernel on the next restart. **But** if a new kernel
hangs, the box keeps trying it. Someone must then switch it off and on. No hardware watchdog is used.
- About 40 packages from Proxmox have ordinary names (for example the disk system ZFS and the Secure Boot loader). So
"which lane" must follow where a package comes from, not its name.
- The internet tunnel program on every box (cloudflared) is 4 months old. Nothing updates it. New row.
- **Thank you for being near the box.** demo-hp restarted twice. It now runs the old kernel and starts the new one at
its next restart.
- **Rows:** 2 closed, 5 opened. The list went from 328 to 331.
## Today (2026-10-04): off-site safety finished
- **The weekly clean-up is ON and no longer stops itself.** The fake-backup check now skips young copies that a manual
backup replaced the same day, instead of refusing. I ran one window on demo-hp: no refusal, nothing removed
(127 → 127), because every candidate was still young. The first real removals come when those copies are older than
8 days — around 11 October. demo-felhom gets its first window at its next night run.
- **DooPlex keeps 8 weekly copies of ep0** (your choice). Old copies go weekly; anything ep0 deletes still never reaches
the copy by itself. Today's nightly copy ran fine.
- **tester-1's 3 old keys are gone.** The daily alarm for tester-1 stopped. (You got one last alarm mail this morning,
sent just before I removed them.)
- **A restore from the DooPlex copy works.** I restored demo-hp's whole box onto a scratch machine in about 3 minutes,
read its data, then deleted it. **One trap found:** a restored box starts automatically and points at the real
drives. Run beside the original, that would be two copies of the same box. I switched it off in time; the runbook
now warns, and a row asks to make it safe by default.
- **"Delete my old set-aside backups" works again.** The hub does it, 7 days after the household's request; the
household or you can cancel in that time. Tested on tester-1 with a planted test folder.
- **Your answers are recorded**, including that you keep the current Hetzner key for now (a row holds the 3 steps).
- **Rows:** 6 closed, 4 opened. The list went from 330 to 328.
## Today (2026-10-03, evening): your choices A and A — built
- **A box can no longer delete its off-site backups.** Both demo boxes now use a key that can only add. I tried a
delete from each box: refused. Old backups all stayed (demo-felhom 11 → 13, demo-hp 91 → 100).
- **No box gets the storage password any more.** The box gives the hub only its public key; the hub puts it in the
storage account. I asked for the password with demo-hp's own login: refused.
- **The hub keeps the passwords locked (encrypted).** A copy of the hub database no longer reveals them.
- **Every day the hub checks each storage account's key file.** It found old unlocked keys from earlier boxes:
4 on demo-felhom's account and 5 on demo-hp's — now removed. tester-1's account still has 3 (its box is gone); you
get a daily alarm for it until they go.
- **The weekly clean-up window works, but it is switched OFF.** I opened one window on demo-hp by hand: it opened,
the fake-backup check refused, and it closed in 3 seconds. The check refused because I had made a manual backup
today — it is too strict. I fix that next; until then nothing deletes old backups (there is plenty of room).
- **ep0's whole-box backups are copied to DooPlex every night.** First copy: 12 GB in 3 minutes, all 4 backups.
DooPlex cannot read them (encrypted per household). A failed copy mails you; the test mail arrived.
- **One bug, found and fixed live:** the first new box version misread its backup count as 0, and you got one
false alarm mail ("demo-felhom: fell from 11 to 0"). **Ignore that mail.** Fixed 15 minutes later (0.289.1).
- **Rows:** 3 closed, 6 opened, 1 opened and closed the same day. The list went from 327 to 330.
## Today (2026-10-03, later): off-site backup safety, step 1 — measured, nothing built
- **The "add only" lock works.** I tested it on tester-1's storage account (your choice; no box uses it now).
New backups go in. Restore works. Every delete is refused. I put the account back exactly as it was.
- **But the lock alone does not protect us yet.** A broken-into box can ask the hub for the storage password.
With that password it can log in and remove the lock. First the box must stop getting the password.
- **A second trap:** with "add only", an attacker can add fake backups dated in the future. The normal
clean-up rule then deletes all the real backups. Any clean-up must check for this.
- **The hub keeps every storage password in plain form.** Anyone who reads the hub database can delete
every household's off-site backups. New row.
- **ep0:** Hetzner cannot snapshot the extra disk at all. One of the three old ideas does not exist.
- **Rows:** 2 closed, 3 opened. The list went from 326 to 327 rows. The dated check for 6 October is done.
## Today (2026-10-03): the to-do list is in order — paperwork only, no machine touched
- **Finished items left the open list.** It went from 442 rows to 326. Nothing was deleted; each moved row names
where its full text is.
- **Every open row now has one category and one severity** (P1 now · P2 before the first paying customer · P3
during the first customers · P4 later). **No row is P1.** 27 rows are P2.
- **Two new automatic checks** refuse a finished row left in the open list, and a new row without a category or a
severity.
- **Your four new items are on the roadmap:** security updates for the box's own system; legal pages and business
papers; "what if the household leaves Felhom"; a second login step for the dashboard. Two of them are also real
findings today: **a box never receives system security updates**, and **the website has no privacy notice,
terms or imprint.**
- The ranked list and my reasoning: the triage recommendation in the audits folder.
**Tester-2 — read only, from the hub.** The customer record exists. Tester-2's box has not registered yet.
## Today (afternoon): every app checked again
- **No data-loss fault.** I checked all 58 apps with the fixed check. It made each app really save something, then
looked where the data landed. **No app saves data where the backup does not copy it.**
- **40 apps: proven correct. 18 apps: not proven either way.** In those 18 the check could not fill every folder (for
example an upload folder stays empty because my test saves no upload). Nothing was found outside a backed-up folder.
I list them and work on them later.
- **papra was broken for new installs, now fixed.** It needs more memory than its limit, so a fresh install crashed in a
loop. I measured it and raised the limit. No box runs papra.
- **plant-it's program image is gone from Docker Hub.** It is already hidden from new installs. No box runs it.
## Removing an app now tells the truth (your choice A)
- When the household removes an app, its own files (books, videos) **stay**. The dialog and the result now say so, and
name the folder. Proven on the scratch box: the video was still there after the remove.
## Before the first paying customer
Everything here must be done before the first customer who pays:
1. **SparkyFitness:** written permission from the author (the e-mail draft is ready). If none → hidden from new installs.
2. **Tandoor:** written permission from the authors. If none → hidden from new installs.
3. **A lawyer reviews the licence list** (Tandoor, SparkyFitness, Emby, n8n, Plex, the paid "enterprise" parts, Redis).
4. **Agent updates:** I may sign agent updates only until the first paying customer; after that you sign them.
Your licence decisions are recorded: Emby, Plex and n8n stay. recipe-importer needs nothing unless we share it.
## Also today
- **New version 0.288.0 and a new golden (0.288.0).**
- **Rows.** 7 closed, 6 opened. The list went from 436 to 442 rows.
## What needs you
0. **The undo choice at the top of today's section** (A: keep the backup as the undo; B: build our own snapshot).
If you say nothing, A stays. Earlier today's two decisions are built. The Hetzner key change stays your call.
1. **plant-it:** keep the hidden template as it is, or remove it entirely (its image no longer exists). **If you say
nothing:** it stays hidden; nothing runs it.
2. **Send the SparkyFitness request, and ask the Tandoor authors** (the "Before the first paying customer" list).
3. **Phone test (2 minutes), only if you want it:** say so, and I put MeTube back on demo-hp with a family login.
## Standing steps
- **Monthly security re-test: last run 2026-10-01, next due ~2026-11-01.** (You start it with the standing brief.)
- **Weekly:** the golden bake (next around 9 October; today's bake was 0.288.0).
- **Registry clean-up:** only when the registry disk fills; "show me" mode first, then a person decides.