Files
felhom.eu/STATUS.md
T
admin 63e0ac01f2
gates / gates (push) Successful in 7s
R-212 CLOSED: the three orphaned stores deleted after a corrected list (~1.45 GB)
The register said 'two set-aside stores, ~1.2 GB'. Measured before touching
anything: THREE set-aside stores totalling ~1.45 GB, and the thing that was
exactly 1.2 GB was demo-felhom's LIVE felhom-repo. Matching on the size would
have deleted a working repository. The operator was shown the corrected list
and confirmed 'delete all three'.

Deleted: demo-felhom orphaned-20260717 (1.4 G) + orphaned-20260718 (3.0 M);
demo-hp orphaned-20260804 (43 M). Both LIVE repos untouched, confirmed by full
listings before and after on each account.

Proof nothing live was caught: a real off-site run on demo-hp immediately
afterwards returned status ok, orphaned false, no error, 6 snapshots.

Method note recorded for the next session: the storage box has a RESTRICTED
shell. 'test -d X && rm -rf -- X' returns 'Command not found' and does nothing
(it failed CLOSED, verified by an unchanged listing); 'rm -r <path>' as one
simple command is the working form.
2026-08-05 11:15:24 +02:00

10 KiB
Raw Blame History

STATUS — what works, what's broken, what's next

Updated 2026-08-05.

A view, not a source. documentation/backlog/OPEN-ITEMS.md is the authority on open work; this page restates part of it in plain words, and nothing may exist only here. Not CONTEXT.md, which is technical state written for Claude Code — keep the two separate. Maintenance: update at the end of every session in which something shipped, broke, or was decided. One screen; cut items rather than extend it.

What works right now

A blank machine boots the Felhom disc, installs itself unattended, and is claimed by the customer, who sets their own password. They install apps from a catalogue of fifty-three, share files over the home network, and open apps from a launcher or a shared link. Backups run on their own to three places — the machine's drive, a second drive, and an encrypted off-site copy.

And the whole backup promise is now proved. On 4 August we destroyed a machine on purpose and deleted a marked file from its disk. Using the recovery code you saved: the backup key came back identical, character for character; the existing off-site store opened rather than starting over; and the file was restored byte for byte identical. (R-201)

What's broken

  • All four recovery steps are now automatic — but a customer still would not know to start. The four steps that stood between "the key is recoverable" and "the file is back" are closed (R-204): the reset-code tool works first time; re-issuing the storage credential no longer falsely marks the recovery key "stale"; the everyday restore says in plain Hungarian that it returned the app's settings and database and not your documents; and, as of today, a rebuilt machine asks for its storage credential itself and the hub answers — no Re-issue click. What is missing is the offer: there is no screen that meets the owner of a rebuilt machine, tells them a sealed package is waiting, takes their recovery code and shows what would come back. So nothing needs you any more, but it still needs someone who knows to look. And the whole journey has not been re-run end to end since these fixes — the four are proved one at a time, not as a single walk. (R-193)
  • Rebuilding a machine still throws away its off-site backup HISTORY. The machine invents the key that encrypts its own off-site backups, and a rebuilt machine invents a brand-new one. Both demo machines did this on 34 August — 51 backups (~1.2 GB) between them. The old key now survives the recovery ceremony, and a changed key now raises an alarm the same day, but the rebuild itself still starts a fresh history. (R-193)
  • One screen still tells the customer something we cannot yet promise. The „elárvult tároló" card says the old backups may later be restorable with the matching recovery code. That is true for machines that re-seal from now on and false for anything already orphaned — and the machine cannot tell which case it is in. We deliberately did not patch the sentence: a conditional promise that can still be wrong is worse there than a vague one. (R-202)
  • The off-site copy can be erased by the machine that made it. The credential that writes it can also delete it. A daily snapshot is armed as a stopgap. (R-95, R-87)

What shipped recently

  • 2026-08-05The orphaned backups are deleted — and the list you were given was wrong, which is why you were asked again. You had approved "about 1.2 GB in two set-aside stores". Measured before touching anything: there were three set-aside stores totalling ~1.45 GB — and the thing that was exactly 1.2 GB was demo-felhom's live store. Matching on the size would have deleted a working backup. With the corrected list confirmed, all three were removed and both live stores left alone; a real off-site backup ran successfully straight afterwards to prove nothing working had been caught. (R-212)

  • 2026-08-05A rebuilt machine now asks for its storage credential, and the hub gives it back. The machine says plainly what it needs — it can tell it has been rebuilt, because its data area is empty and the hub is holding a sealed recovery package for it — instead of leaving the hub to guess from a silence that has four possible meanings. The hub waits long enough to be sure it is not a restart, then re-uses the credential it already holds before creating a new one at the storage provider. The recovery ceremony stays manual, deliberately: a credential can be replaced, your recovery code cannot. (R-204 item 4, R-193)

  • 2026-08-05The DooPlex server's disk is out of danger: 86% full → 54%, and the storage layer is unstuck. The cause was leftover working data from building our own software — 157 GB of it, growing about 5 GB a day, which nothing was allowed to delete. 148 GB came back in 86 seconds. The storage layer had already stopped accepting new copies of any volume onto that disk; that is fixed the same day. A 30 GB ceiling is now in place and was proved to work by deliberately overfilling it and watching it evict — not by assuming the setting took. Two things that failed quietly around it: the warning meant to catch exactly this could never fire (fixed and proved), and the weekly cleanup is still forbidden from touching the thing that grows (next session). (R-205 … R-211)

  • 2026-08-05The build data now lives on the second SSD, moved with nothing lost. All 345 images, every saved volume and both development databases came through identical — checked before the original was touched and again afterwards, and confirmed by running a real build on the moved copy. The server's own services never went down: Gitea, the registry, the hub and the backup system run on a separate system and stayed up throughout. The second SSD now also reserves 80 GB for this, so the storage layer can no longer quietly claim the space and repeat what happened to the first disk. The old copy is kept as the way back until the machine next restarts. (R-209)

  • 2026-08-05Three of the four recovery crutches removed. The reset code works first time; a credential re-issue no longer blocks off-site backups on a healthy machine; and the default restore no longer quietly returns the wrong thing. (R-204 items 13, R-196)

  • 2026-08-04 (night)The drill PASSED, end to end, on real hardware. (R-201)

  • 2026-08-04 — The folder-left-out-of-the-backup problem fixed both halves: the app and its backup look in the same directory, and a backup that misses a folder marked essential reports incomplete instead of success. (R-203)

  • 2026-08-04 — The hub now keeps the off-site backup key when a machine re-seals, instead of only the whole-machine one, and a machine can fetch its own sealed package back. (R-198, R-199)

  • 2026-08-04 — The daily false alarm about David is gone; a vanished permission now repairs itself and says it had to; the weekly off-site backup stopped reporting failure after a successful upload. (R-195, R-190, R-191)

What we're working on

  • Next: the retention proof. The hub keeping the old sealed key when a machine re-seals is the one remaining link that has never run outside a test. Proving it needs a second deliberate wipe on the spare demo machine, and it is its own procedure. (R-198)
  • Then: the orphaned-backup deletion you asked for (below), and the off-site copy the machine can still erase. (R-193, R-95)

Waiting on you

  • One thing to read after the machine next restarts — and nothing to do until then. You told me not to restart DooPlex, so I did not, and the move to the second SSD has therefore never been through a restart. It works right now and nothing was lost, but a restart is the one test that matters for this kind of change, and it has not happened. I made it check itself: whenever the machine next starts, for any reason, it writes a plain PASS or FAIL line to /var/log/felhom-store-postboot-check.log. If it says PASS, the old copy can be deleted and 34 GB comes back. Until then I have deliberately kept that old copy, which is the only quick way back — it is why the disk sits at 54% rather than lower. (R-209a)
  • One list to rule on: 193 old images that exist only on this machine. 131 controller versions and 62 hub versions are not in the registry, so they cannot be re-downloaded — all of them old (controller up to 0.135.0, hub up to 0.57.0; everything newer is safely in the registry). Nothing was deleted. Worth knowing before you spend time on it: they only account for about 27 GB against 199 GB now free, so this is about clutter, not space. (R-210)
  • Nothing.
  • The recovery screen you described has been priced, and it can be built. A freshly installed machine that finds a sealed package waiting should say so, offer a box for the recovery code, and show what would come back before doing anything. One thing to weigh, deliberately not decided: that screen is reachable by anyone with the household's dashboard password, and the preview reveals backup dates and app names. (R-193)
  • The orphaned backups on the storage box — you said delete, and it is still owed. About 1.2 GB across the two demo machines, in set-aside stores nobody can open and nothing prunes. It wants its own session rather than riding along with other work. (R-193)
  • (decided 4 Aug) You chose not to keep a copy of the backup key on the Proxmox host, which makes the customer's own recovery code the only route back from a rebuild. (R-193)
  • A job, not a decision: the hub password needs changing. A diagnostic command printed it into a session log; nothing suggests anyone else saw it. (R-132)
  • One small question, not urgent. The automatic version check cannot see which version you have told machines to install, only which ones exist. (R-184)