R-449. Until today one app upgrade out of 53 had ever been measured, by hand, and the whole update arc was designed against that single data point. C3 first: the negative control, whose TO image exits immediately, came back failed. That is what makes the greens mean anything, and it cost 556s because a negative is only honest if it waits out the full settle window. Seven edges, three apps. All five real catalog upgrades kept the customer's data. The finding that changes an assumption the arc was carrying: whether an upgrade can be UNDONE is a property of the individual APP, not of upgrades. Docmost refuses - 'corrupted migrations: previously executed migration 20260213T085259-notifications is missing' - and privatebin does not. That reproduces the Nextcloud result on a second app by a DIFFERENT mechanism, so the struck word 'rollback' now rests on two measurements instead of one. The finding nobody was looking for, R-459: our own bookstack template moves MariaDB across a major and sets no MARIADB_* env at all, so the engine logs that the datadir upgrade it requires is being skipped, and serves anyway. The cause is assigned rather than guessed - the app half alone produces no upgrade line, both edges that move the engine produce it - which is exactly what decomposing E3 into E3a and E3b was for. It also explains why E3's abort looked like it worked: the datadir was never converted. Whether that ever breaks is NOT established, and the row says so. Also opened: R-460 (bookstack's file half cannot be seeded headlessly), R-461 (target-selection.md names a venue that does not exist and fences a VM that is gone), R-462 (the widening, costed with this run's real numbers - and the cost is dominated by fixtures, which do not amortise). Teardown all three layers, hub checked rather than asserted. local-lvm read 30.50 percent before and after. The capability map was deliberately NOT edited: this measured apps, not the product.
31 KiB
STATUS — what works, what's broken, what's next
Updated 2026-09-06 (second pass) — I built a machine that upgrades a real app with real data in it and then asks the app whether the data is still there. Three apps, five real upgrades: the data survived every time. It also found a genuine problem in our own BookStack setup. ONE NEW THING NEEDS YOU: item 11 — how wide should I take this?
Earlier 2026-09-06 — a restart no longer changes which version an app runs. Fixes still arrive every 15 minutes, and a broken app definition still repairs itself. Only the Update button moves a version now. Live on the HP (0.235.0). Two old items closed: the Hetzner e-mails are answered, and the Docker Hub login is in place.
Earlier 2026-09-03 — you spotted that OpenGist had no label. You were right, and it was a real gap: the label only appeared on apps something had restarted. Fixed and live (0.234.0). Every app on both machines now carries one.
Earlier 2026-09-02 — the box now writes down which version of each app it is running, and shows one small label saying whether it is up to date: „Naprakész" or „Frissítés elérhető — 52 napja". No version numbers, and nothing about updating changed.
Earlier 2026-09-01 (third pass) — the backup work is FINISHED for beta, and I have written down where it stops. One alarm that was telling you something untrue is fixed and live (hub 0.111.1). ONE THING is waiting on you: two short e-mails to Hetzner, drafted and ready to paste.
Earlier 2026-08-31 — I measured whether the box could test its own remote restore without you, then built the narrow version you picked. The measurement is why it is 3 seconds a night and not an evening's work.
A view, not a source.
documentation/backlog/OPEN-ITEMS.mdis the authority; this page restates part of it in plain words, and nothing may exist only here. Items, not paragraphs. One screen. If it does not fit, it belongs in the register instead.
Waiting on you
This section is allowed to be longer than one screen, and each item says what happens if you do nothing.
-
One thing is waiting on you: item 11 (how wide to take the upgrade testing). Item 4 (the Hetzner e-mails) is answered and is being handled in a separate session. Item 7 — the safety-copy decision — is the one open question, and it is not urgent any more: the thing that made it urgent was that a restart could upgrade an app behind your back, and as of today it cannot. Item 10 is new and needs nothing from you. Item 5's alarm mail can now be ignored for good. Otherwise: Both problems the overnight test found are fixed and proven on the real machines:
- the background job that could delete a live restore's lock now waits its turn — and the check that finds the next one like it is a test, not a comment, so it cannot come back quietly;
- the nightly backup check now runs on
demo-felhom. It said pass there this morning, on a machine where it could not start at all yesterday. The golden carrying both is baked, vouched, and the floor is raised.demo-felhompicked it up by itself in about 20 seconds.
-
Whether to keep the test app
bentopdfondemo-hp. I deployed it last night because it is the ONLY app of our 53 with neither a database nor stored files — which makes it the only way to prove the new check does not cry wolf on an app that legitimately has nothing. It passed silently, which is what we needed to see. My pick: keep it, as a permanent control. If you do nothing: it stays, using almost no space. Say the word and I remove it. -
Nothing else about this release. Everything in 0.232.0 ships in the controller image plus two register lines in the hub (already live). No customer action, no data migration, no credential change.
-
Please send two short e-mails to Hetzner.DONE — you sent them and Hetzner replied (2026-09-06). The reply is being worked in a separate session; nothing about it belongs to the update work. The background below is kept because it is why the questions were asked.felhom.eu/documentation/runbooks/provider-questions-2026-09-01.md— open it, copy, send. No password or key is in that file, and none should be added.Why. Yesterday I told you the snapshots make a wiped backup survivable: lose about a day, copy the rest back file by file. The first half is still true. The second half is not, and I found that out by trying it. I tried 777,600 snapshot names on the storage, over nine days, in Hetzner's own naming style. None of them opened. Then I found why: your data and the snapshot door sit on two different drives inside the storage, and the door for your data does not exist at all. So there is no way in from the machines.
What is still true, and it matters: a machine that wipes its own backup still cannot touch the snapshots of it. The older copy is there. What we do not have is a way to reach it.
The two questions. One: can the main account pull single files out of a snapshot? Two: on one of Hetzner's own tools, is a "cannot delete" switch forced by them, or chosen by the machine? The second one could remove the whole problem — no new hardware, no moving anyone's data.
If you do nothing: we cannot finish this. The backups keep working and keep being checked; we simply cannot say what a wiped backup costs, and I would then put this risk back near the top of your list. My pick: send both. It is five minutes and it decides an evening's work.
-
You will have received an alarm email from me today about
demo-hplosing 65 backups. It is a test and nothing is wrong. I built the new "someone deleted the backups" alarm and had to fire it once for real to prove it reaches you. Subject:[Felhom] 🔴 demo-hp: offsite_snapshots_dropped. No backups were deleted. If you do nothing: nothing — but please do not act on that one mail. Closed now: that was the only such mail, it was a test, and the alarm's wording has since been corrected (item under Decided below). Nothing further is needed from you here. -
Whether to change the hub password (R-350). I printed it into my own session log on 20 August. Not in git, not in any saved file — in the log on this machine. If you do nothing: it stays as it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
-
Where should the safety go before an app updates? This is the one decision from today's measurement, and it is a design choice, not a bug report.
What I measured. The box downloads new app versions by itself every 15 minutes and writes them into the customer's files, whether the app is running or not. Nothing tells the customer. Then the Restart button — not just Update — installs that new version. I watched it download a version that was not on the machine and swap the app onto it in 18 seconds. And the box does it on its own when an app fails to come back after a crash: nobody pressed anything.
One fear is smaller than we thought, and you should have that too. A plain power cut does not upgrade anything. The apps come back on their old version. It only happens when an app fails to return.
One fear is bigger. I tested whether we can undo an app update. We cannot. Once an app has moved its data to the new version, putting the old version back gives an app that will not start at all. So "rollback" is the wrong word and I have struck it. The only way back is to restore the customer's data from a copy taken before the update — and today no update takes one.
The decision, in one sentence: should the safety copy sit under the Update button only, or under everything that can install a new version?
- Under the button only. Cheap and quick. Covers the case a customer causes. Leaves the unattended path uncovered — the one where an app that failed to come back is upgraded with nobody watching.
- Under everything. Covers all of it. Costs more, and it has a hard limit I measured: for a big app a copy is roughly 30 minutes and about twice the app's size, against a standard box that ships with 20 GB. A copy of a large app does not fit. So this option cannot be built without also answering where the copy lives.
My pick: under everything — but decide the "where does it live" question first, because the answer decides whether the rest is even buildable.
If you do nothing: nothing breaks today, and no customer is at risk this week — the fleet is young and its running versions match the catalog. But the exposure is real and dated: Peti's box has a one-major upgrade of
ralllyqueued behind its next boot, from a catalog change made three days after it went quiet. Who else is blocked: nobody can spec the safe-update work until this is answered, because the two options produce different products.Full measurement, with the controls and the quoted output:
felhom.eu/documentation/audits/SPIKE-app-update-2026-09-01.md. -
Nothing here needs you. Two things I got wrong, both now fixed.
You found the first one. OpenGist showed no label. That was not a misunderstanding — the label only appeared on an app that something had restarted, so an app that simply runs showed nothing, possibly for months. On a quiet machine that is every app, which is exactly the machine we most want to be able to look at. Now the box reads what every app is on when it starts up, and writes it down. It only looks — it starts nothing and changes no app. Live on both machines: all nine apps on the HP now carry a label, and so does OpenGist.
One thing to expect, so it does not look broken: the label says „Naprakész" when an app is on the newest version, with no number at all. A number only appears when the app is behind — and then it is how long the newer version has been waiting, not how old the running one is. That was your ruling and I think it is still the right one.
The second one was mine and smaller: a test I wrote pinned a date and an age, so it passed the day I wrote it and failed the next morning. Fixed, and I have written down that the same trap may sit in six other test files — named, not accused; someone has to read them.
The box now writes down which version of each app it is really running, and shows the customer one small label: „Naprakész" or „Frissítés elérhető — 52 napja". No version numbers — a household cannot act on
26.05.2.Both halves are proven on the real machine. I made two apps fail to come back, the box repaired them itself and wrote down exactly what it installed — one line per container, with the fingerprint that cannot lie. I checked those fingerprints against the machine independently and they match. Then I opened the real customer pages and read the labels off them.
Where I went wrong, because you should not have to spot it twice. I told you the saved password no longer worked on either machine. It worked fine. I read it out of the file wrongly — the value is wrapped in quote marks and I only removed one kind. You caught it in one line. The machine told me "wrong password", which was true, and I took it to mean the password was wrong when it meant what I sent was wrong. I have filed that as a small job for myself: one shared way of reading that file, so the next session cannot get it half right.
If you do nothing: nothing. This one is closed.
-
Nothing here needs you. The version of an app is now frozen, and only the Update button moves it.
What used to happen. The box downloaded the catalog every 15 minutes and wrote each app's new definition straight over yours — including a new version. From then on, thirteen different things could install that new version: the Update button, the Restart button, or one of eleven repairs the box performs on its own. Nobody had to press anything.
What happens now. The version is pinned to what you have. Only a deliberate Update moves it.
What deliberately did NOT change, because it was worth keeping. Corrections to an app's definition — a fixed health check, a memory limit, a new setting — still arrive on the same 15-minute cycle, and a broken definition still repairs itself. That was a real benefit of the old behaviour and it is intact. In one sentence: while the catalog offers the same version you are running, its fixes reach you; the moment it moves to a newer version, you stay where you are until you choose to update.
Proven on the real machine, twice over: I pushed a genuine catalog change with no version in it and watched it arrive on the normal cycle; then I pushed a version change and watched the app refuse it.
Two things this did NOT do, so they are not read as done. The Update button is exactly as safe as it was yesterday — no backup, no undo. That is the next piece of work, and it is item 7. And an app the box could not confidently pin would have been left behaving exactly as before, loudly; on the HP that was none of the nine.
One rough edge I chose to write down rather than fix: a frozen app still receives the small metadata file that carries its health check, so it can be given a check written for a newer version and look unwell when it is fine. It cannot lose data — the worst case is a false alarm. Freezing that file too would break the „Frissítés elérhető" label, which is a worse trade.
-
How wide should I take the upgrade testing? This is the one decision from today, and it is about money and time, not about safety.
What I built. A machine that installs an app, puts real data in through the app's own front door, upgrades it, and then asks the app for the data back. Not "did it start" — the box has already fooled us that way once.
I also taught it to fail. Before believing anything, I pointed it at an upgrade I knew was broken. It came back red. That is why I trust the greens.
What it found, on three apps and five real upgrades.
- The data survived every single time. That is the good news and it is worth having.
- Whether an upgrade can be UNDONE depends on the app, not on upgrades. Docmost will not go back — the old version refuses to start on the changed data. PrivateBin goes back fine. We had been assuming one answer for all 53 apps. There isn't one.
- And it found a real problem in our own BookStack setup. Our template moves the database engine to a new major version, and the engine says, in its own words, that the conversion it needs is being skipped. It works today. I did not measure whether it ever breaks, and I am not going to guess.
The decision: how many of the 53 apps do I test?
- Only the apps that keep data in a database (~25). Today's run showed the undo question only ever bites there. Cost: roughly a day of my time.
- All 53. Complete, and it gives us a list nobody has. Cost: several days, and the reason is not the machines — it is that each app needs its own hand-written way in, and that work does not get cheaper the more you do. Two of today's three needed one, and one of them took two attempts.
My pick: the database ones first. It answers the question that changes the product, and if it goes well the rest is a decision you can take later with better numbers.
If you do nothing: nothing breaks. The machine is built and committed, so it does not go stale, and the BookStack problem is written down and waiting either way.
-
demo-hp's network setup does not match our own notes (R-338) — the machine works, the page is wrong, or the other way round. If you do nothing: the page keeps misleading the next session, as it misled one by an hour.
Decided — and what would reopen each
-
THE BACKUP AND RESTORE WORK IS FINISHED FOR BETA. DECIDED 2026-09-01. What is done, and proven on the real machines: everything you or a customer does alone — getting deleted files back, getting an app's data back, getting a whole app back, and losing a drive. The restore tells you what it put back, refuses if there is no room, will not accept a half-copy, and puts your own data back if it fails. What is parked until after beta: everything only I do, with you — rebuilding a machine as itself, losing a whole box, recovering from ransomware, restoring the hub, and losing Hetzner. Six of these have never been timed, and the hub has never been restored. They are written down, they are real, and none of them stops a beta customer. Reopens if: something a customer does for themselves turns out to be broken; or Hetzner's two answers change what the snapshots are worth; or a real customer's data is at stake in one of the parked items.
-
The alarm that promised too much: FIXED and live (hub 0.111.1). DECIDED 2026-09-01. Yesterday's alarm mail said a deleted backup was "recoverable file-by-file". We now know it is not. I removed the promise rather than writing a new one, so the sentence stays true whatever Hetzner answers. It no longer says the data is lost either — that is still usually untrue. Reopens if: Hetzner's answers give us a real route back; then the alarm can name it.
-
The missing-golden warning is now pointed at the people who can act on it (R-404). DECIDED 2026-09-01, and built the same day. The warning was aimed at the wrong repository: the one where a release actually happens never checked at all, while the one that only holds documents was refused on every push — including the push that RECORDS a golden bake, which is the very act that clears the warning. So the check was blocking its own cure, and we had skipped it thirteen times. What changed: the release repository now prints a reminder the moment a release is committed (it never blocks — you cannot bake a golden for a version you have not pushed yet), and the documents repository still runs the check on every push and still says so loudly, but only refuses a push that touches real code. Every other check still blocks everything, always. Nothing was silenced and no product code changed. Reopens if: a release ever ships without a golden and nobody noticed — that would mean the reminder is not reaching anyone, and the answer would be to make the release repository refuse rather than remind.
-
Getting old backups back yourself: NOT BUILT, deliberately. Reopens if: a real customer asks. (R-312)
-
The unopenable old copy on
demo-felhom: KEPT as a test fixture — the only state in existence where a set-aside store is present and cannot be opened. Delete when: that work ships or is abandoned. (R-313) -
A machine in two kinds of trouble says both things: LEFT AS IT IS. Reopens if: observed outside a constructed test. (R-303)
What works
Both demo machines are home, healthy and reporting — agent 0.130.0 published and running on both.
Which controller each box runs, and where the floor sits, is item 1 above and is not restated here —
R-395: this paragraph carried a second copy of those numbers, it went stale by seven releases, and the
page then disagreed with itself about the thing an operator checks first. Ask the hub (/hosts,
/configs) or the box for what is live; a doc is never the authority on a version.
Off-site is credentialed on demo-hp and its store opens with the machine's own key.
The fleet, because two summaries have been misread: five customer records, three machines.
demo-felhom and demo-hp are ours and disposable; drill-r50 is a nested drill VM, reverted and
off. peti-felhom is a real machine we have not heard from since 15 July and has no host record.
tester-1 is a record with no machine.
Shipped
-
Something finally checks that the off-site copies are still there and readable (R-359 + R-397, controller 0.227.1, proven on
demo-hp). Until today nothing did — not the box, not the agent. The whole-machine backups had their own checks; the copies holding your customers' documents and photos had none, so we would have found a problem at restore time, with a customer waiting. Now the box checks its own off-site store about once a week and tells you only if something is wrong. A pass sends no e-mail, on purpose — a weekly "everything is fine" is how people stop reading their alerts. It also catches itself up: it asks „has it been more than seven days?", not „is it Sunday?", so a machine that was switched off on its check day is checked the next day. And it never gets in the backup's way — if a backup or restore is running, the check steps aside and tries again tomorrow. That is not politeness: the tool would otherwise clear a lock that a live backup was holding. Proven live, twice over — a real check against your real store (35 seconds), and a second check fired during the first, which correctly stepped aside without doing anything. Read item 2 under „Waiting on you" for what this check does NOT see. -
The product stopped claiming a check it never ran. The monitoring page said an integrity check ran every Sunday. It did not exist. The e-mail text, the settings checkbox, the hub's side of it and a debug button were all built and wired to nothing (R-397). They now have the missing piece.
-
Taking the safety copy no longer destroys the app's own backup (R-361, controller 0.221.1, proven on
demo-hp). Before every restore the machine saves a copy of your live database. To do that it called the ordinary backup routine — which always writes to the app's normal backup filename first — so the app's real backup was overwritten and then renamed away. Until the next nightly run the app had no database backup of its own, and a local recovery in that window would have told you the app never had a database. A comment in the code said this could not happen; it could, and had been happening for four months. Proven fixed the only way it can be: the app's own backup file is now byte-identical before and after a restore, on both database types. -
A failed database restore now puts your data back by itself (R-379/R-380, controller 0.220.2, proven on
demo-hp). Until today, if a restore of an app's database went wrong, the machine had already taken a copy of your live database — a good copy — and nothing in the product could put it back. You were shown a filename. On one of the two database types it was worse: part of the restore applied, part did not, and the dashboard said the app was healthy. Now the machine puts your own copy back automatically and says plainly: the restore failed, your data is as it was, the app is running. Proven on both database types, byte-identical both times. If even that fails, the app is deliberately stopped and held rather than started — a running app on a half-written database lets you type into it and makes the damage permanent — and you are told to contact us. That was your ruling this morning. Two things also stopped: the error no longer pastes raw database text at you (it was 615 bytes once, including rows out of your own database), and the undo copies no longer pile up forever — three per app, and they were being copied off-site permanently. -
An empty package can no longer wipe out a good one (R-403, controller 0.230.0). Yesterday we wrote this down as suspected and said plainly it had not been tested. We tested it first, and it was real. On a demo machine, on yesterday's build: an app's copy on the second drive went from 120 MB — four database backups and three data archives — to 7 KB, nothing left, in a single nightly run, and the run reported success. The cause was that the nightly job only asked does the folder exist before copying over it, and an empty package is a folder that exists. Now the nightly job refuses to replace a complete package with an empty one. It keeps what it has, says so on the app's own backup page, and carries on with everything else. Proven on the same machine, in the same state: all seven files still there, byte for byte. Two more things came with it. The page no longer calls that copy fresh when the run did not refresh it — it names the real date of the package instead. And after a restore from the second drive, the first drive's package is filled back in immediately, so the empty state that started all this cannot happen again. A copy that legitimately gets smaller still gets smaller; only the one dangerous case is fenced.
-
The copy on the second drive can now bring an app back (R-102 + R-103, controller 0.229.0, proven on
demo-hp, and delivered — golden 0.229.0 is vouched and the fleet floor is raised, so a machine installed today has it). Every night the box copied each app's whole recovery package onto the second drive — its settings, its database and its data. It did that for months. Nothing could open those copies. No button, no screen, no command. That mattered most in the one fault the second drive exists for: if the first drive dies, the package on it dies too, and the copy that survived could not be read. For 45 of the 53 apps that is everything they own. Now the same restore that always worked from the first drive can read the copy on the second one, and the button is on the app's own backup row. Proved with the first drive's package taken away: Docmost came back in 29 seconds — its database, all three data volumes, a file with a Hungarian accented name back byte for byte, and the app then read its own rows with its own password. Done again with the app's password file also taken away: the copy carried the passwords too (2 of 2). The screen that used to say „press that other button on another page" now offers the action itself. It says plainly that this one overwrites what is there — the gentle „Fájlok visszaállítása" beside it still only adds back missing files — and it names the date of the copy, so nobody puts last week over today by accident. -
We counted the apps this affects, and settled it. Two of our own notes disagreed — 43 or 45. The answer is 45, counted with the product's own rule against the live catalogue. The older count missed radarr and sonarr.
-
The off-site restore now works for the other 40 apps (R-356, controller 0.219.0, proven on
demo-hp). It used to refuse before starting, tell the customer a running app „nincs telepítve", and send them to reinstall it "to the same place" — a place those 40 apps never offer, because they were never given a drive to choose. It was asking one question to answer two. Proven today onprivatebin: data planted through the app itself, backed up, deleted, restored — all 15 files back byte for byte, Hungarian accented names included, message „0 fájl és 1 adatkötet visszaállítva". The 13 apps that do have a drive are unchanged, checked the same way. -
The off-site restore gives an app's data back at all (R-354, controller 0.218.0). It used to say „0 fájl visszaállítva", report success, and the folder was simply not there. It now names what came back, because a restore that mentions only its file count is how a silent loss reads as a success.
-
Paperless's database is in the backup, and restoring it takes an undo copy first (R-355, controller 0.218.0). The dump was landing in a folder named after an app that does not exist. One app of 53 was affected, established with a check first proved able to catch a planted second case.
-
The system tells you when it cannot see the off-site copies (R-339) — a mail after ~30 minutes, hourly while it lasts, one all-clear. Caveat: it watches whether the machine answers, so it would not have caught the 18 August fault, where one service was wedged and the machine stayed healthy.
-
The connection leak was ours and is fixed (R-344). Our agent opened a connection to the off-site box every 15 minutes and never closed it; the idle timer was switched off. Proved by fixing one machine and leaving the other: same work, 4 more leaked on the untouched one, none on the fixed one. The off-site box is back to 17 open connections from 415.
-
A dated check can no longer be quietly missed (R-341) — but it speaks on the next push, not on the day. A machine we tell to be quiet is no longer reported as dead (R-321). One name per secret (R-295, R-323). The hub's own words are under a guard (R-324). Removal reverses the installation (R-316). A correct recovery code is no longer called wrong (R-311). The drive can be re-attached after a reinstall (R-280).
Broken, or knowingly incomplete
- We do not know what the deep check costs on a BIG store (R-401). Since 0.228.0 the weekly check re-reads all your stored data, not just the list of it. We had to: a copy was damaged in a way that left its size unchanged, and the old shallow check said „no errors were found". Only the deep check caught it. The cost we measured was four seconds — 35.0 s before, 39.2 s after — but that was on a 134 MB store, and it will not stay four seconds. So the box now tells us: any check that takes longer than five minutes writes a warning naming this item. If you do nothing: every machine re-reads its whole store every week, however large it grows, and the first person to notice would be a customer whose upload is busy. The warning is there so that does not happen.
- We ask the off-site box a question about once a second (R-336) — ~85,000 a day for a box we write to weekly. The leak that made this dangerous is fixed (R-344); the volume is not. The ceiling is under a year away on the corrected measurement, not two.
- Peti's machine has no recovery route at all. A real machine belonging to a real person, silent since 15 July, no key, no off-site copy, no local backup. If that drive fails, everything on it is lost. First act of any visit: copy the ~3.6 GB off before anything is reinstalled.
- The agent picks dnsmasq by looking at a file another package owns (R-317) — one line; LAN name resolution goes missing quietly.
- Three facts the machines send still have no reader (R-264); the storage page has its own
reason for an empty list (R-298); two thirds of the standing picture is unproven (R-326:
23 of 55 claims walked —
python3 scripts/unproven.py); the picture still describes one defect we fixed twice (R-327).
Working on next
The 2026-08-09 batch (R-279 … R-292), still untriaged; the three remaining R-264 readers; R-317 (one line in the agent); R-327 (decide the naming claim's status); R-359 (nothing reads the off-site store).