Files
felhom.eu/STATUS.md
T
2026-08-10 14:29:54 +02:00

15 KiB

STATUS — what works, what's broken, what's next

Updated 2026-08-10 (afternoon).

A view, not a source. documentation/backlog/OPEN-ITEMS.md is the authority; this page restates part of it in plain words, and nothing may exist only here. Not CONTEXT.md, which is technical state written for Claude Code. Items, not paragraphs. One screen. If it does not fit, something belongs in the register instead.

Rebuilt from the register on 2026-08-07, from 258 lines. The old "what shipped recently" log is what the per-repo CHANGELOG.md files and the register are for, and is not restated here.

Controller 0.211.0 is baked, vouched and DELIVERED to fresh installs — the fleet is a separate switch

The tester-visit release shipped: the data drive can be re-attached after a reinstall (the restore page's „két kattintás" was pointing at an empty picker — it was zero clicks); the orphan card stops promising the set-aside off-site copies may be restorable, which the machine showing that card cannot know; and the dashboard code is now called „Beállító kód" everywhere — „Visszaállító kód" is retired, because it collided with the escrow „Helyreállítási kód" and that collision cost a real code.

Golden 0.211.0 is published and vouched (agent 0.128.0, min agent 0.127.0), verified by re-downloading the served bytes and hashing them.

One thing needs you. The global update floor is 0.200.0, and boxes auto-update to the floor, never to the newest. So fresh installs get these fixes and the existing machines do not — demo-hp and demo-felhom stay on 0.210.0 until the floor is raised. That is a one-field change on the same page, and it is deliberately yours.

Two things were dropped and are not forgotten: our own uninstall still leaves dnsmasq holding :53, so the next install refuses and blames the household's network (R-293 area, untouched); and the hub's own emails still call the setup code by the retired name and send people to a page a rebuilt box does not show (R-295, half done).

The installer's stale-golden fix is written but NOT published — an install could silently reuse an old archive lying on the machine, including one too old to run the recovery screen. The fix is in main, which publishes nothing; the tag is deliberately uncut until we have watched the failure happen once on a drill machine (R-297).

Both machines are home, unmuted and healthy — one thing still needs you

Back online 2026-08-10 ~09:26 CEST, both unblocked on the hub, both reporting OK on the approved pair (agent 0.128.0, controller 0.210.0). No false alarm fired on power-up. drill-r50 is untouched and still blocked, as intended.

demo-hp is in good shape. Its off-site repository still opens with the machine's own key — 18 snapshots, including yesterday's rehearsal files — so the tier is credentialed and ready; its first scheduled run since the rebuild is tonight at 04:15. Two apps it had before the rehearsal (opengist, privatebin) were never reinstalled; only Calibre-Web was, as the walk needed.

demo-felhom is protected again — and the recovery you authorised turned out to be impossible. Before running it I checked, and the sealed package holds the same key the machine already had — a key that provably does not open its own backup store. Recovering it would have handed back something useless. The store was written under an older key whose sealed copy was not retained (the retention fix landed hours too late for it), so those 1.2 GB are permanently unreadable by anyone, including us.

So I took your stated fallback: the old store was moved aside, not deleted (/home/felhom-repo.orphaned-20260810), a fresh one was created under the current key, and a real backup ran — succeeded in 10 seconds, and I listed what is inside it rather than trusting the green tick: OpenGist's configuration, its manifest and its data volume. The week without off-site protection is over. (R-278 closed.)

One thing that needs your judgement, not mine. The card that offered this told the customer their set-aside backups may be restorable later with their recovery code. For these ones that is simply untrue, and it is said to precisely the people who have just lost their history. (R-202 — now evidenced.)

Can anyone else lose their history the way demo-felhom did? No.

One read of the hub's own records, no machine touched. The hub holds backup keys for exactly three machines. Both demo boxes lost their old key in the same four-hour window on 4 August, before the retention fix was in force — that is the whole population of the problem, and it is entirely ours. The tester's machine has no record at all, so it cannot be affected; and anything enrolled from now on is covered, because the fix has been in force since 4 August.

I ran a control before trusting the query: it had to say material present for a machine known to have it and absent for one known not to. It did both.

Three green dots came back, nine stayed grey, and one rule finally fired

The nine greys are the honest number. For those, no document anywhere walks the claim, and saying so is more useful than a dot nobody can defend.

  • Back to green, each citing the document that walked it: the drive wizard (a live drive taken through scan → format → mount → enrol), the on-box app backups (an overnight destructive campaign across both machines), and the lost-recovery-code case — which we then proved the hard way this morning.
  • One claim stayed grey for a new reason, and it is the interesting one. The unattended restore-proof does have a receipt from 28 July — but demo-hp's restore-test failed on 5 August and the machine has since been wiped and rebuilt. It is a claim about something that keeps happening, so an old observation cannot carry it. This is the first time that rule has fired; two nights ago it fired zero times out of twelve.

The prune mystery is solved, and the answer was written down all along

Who deleted the old versions: you did, on 4 August evening, on your own rule — 33 deletions, keep set asserted first, every one a clean 204. It was recorded inside the row about the Configuration page being slow, because pruning artifacts is what made that page fast. Two sessions failed to find it. You are no longer blocked on establishing something that was already on file.

I also got a number wrong yesterday and it is corrected: I said the container packages held nineteen versions and used that to argue against the prune. Counted properly — with pages — they hold 270 and 169, and the two that were pruned sit at exactly ten each.

The sentence we should stop saying

When a machine's off-site history is set aside, the card tells the customer it may be restorable later with their recovery code. The machine showing that card cannot know whether it is true — the fact lives on the hub and is not sent to the box. For anything set aside before 4 August it is simply false. The replacement wording is written and waiting (documentation/design/SPEC-orphan-card-copy-2026-08-10.md); it ships with the next controller release so one image bake and one approval cover it, rather than costing you two of each.

What works

A blank machine boots the Felhom disc, installs itself unattended, and is claimed by the customer, who sets their own password. They install apps from a catalogue of fifty-three, share files over the home network, and open apps from a launcher or a shared link. Backups run on their own to three places — the machine's drive, a second drive, and an encrypted off-site copy.

The backup promise is proved, and so is getting the data back yourself. A machine has been destroyed on purpose and its files came back byte for byte identical — four times now. On 2026-08-07 the household's own journey passed for the first time: someone with a browser and their recovery code got everything back with no command line inside the machine at any point, in 72 seconds. The two rough edges that walk found are also gone. (R-201, R-252, R-253 — closed.)

Shipped 2026-08-09 — the guards, and the rehearsal that earned them

An approval that cannot be installed is now refused at the moment you press Save: the hub checks the version's git label and that its file downloads, and refuses with a message naming the fix. A second machine catches it a step earlier in the agent repo. Both were owed after every install in existence failed for hours on 2026-08-09. Hub v0.102.0 is live.

The reinstall rehearsal: a demo machine was wiped and put back. All four test files returned byte for byte, accented Hungarian filenames included, checked as raw bytes. But it only finished because a terminal was available twice — the install died on a missing version label (R-273, now guarded) and a reinstalled machine still cannot re-attach its own data drive (R-280 — the one to fix before the tester's visit). Detail: documentation/audits/REHEARSAL-byo-reinstall-2026-08-09.md.

What's broken

  • Nothing else new is broken. The check against a fourth secret-in-a-page covers 4 pages of 27, and the cheap one covering all of them is blind to the shape that shipped. (R-255)
  • An already-paired box is still told to pair itself, 25 minutes on. (R-214, R-235)
  • A backup that covered nothing still calls itself „Sikeres". (R-240)
  • A machine waiting for its recovery code can stop backing up off-site without alarming us. (R-243)
  • The card offering to reopen set-aside backups promises more than we can deliver. (R-202)
  • Deleting a customer leaves rows behind while reporting a clean teardown — no secrets, but it accumulates. (R-244)
  • Putting restored files back where they belong is still manual. (R-213)

Three rulings, written down so they stop living in a conversation

  • The managed-update floor. It is deliberately parked, and the trigger to raise it is the first machine that is not ours; after that it moves with the publish train. Worth knowing alongside it: the updater always aims at the floor, never at the newest, so a machine at or above the floor updates to nothing. Correction to the number that was going round: the floor is live at 0.200.0, not 0.156.0 — checked twice today, on the hub page and in both machines' own logs.
  • R-264 is decided. Build a reader for guest-network health, the staged-update pair, the restore-test depth pair, and the two backup-integrity timestamps. Record a stated "no reader wanted" for the repaired-recently flag, the tier-applied timestamp, the config fingerprint, the per-stack object and the drive-migration marker. Decide the reporting-disabled flag on its own merits — if that is a state we support, it must be visible or a staleness alarm will one day fire on a machine that is fine. A "no" ends by changing the allowlist reason from arguably owed to deliberately not consumed — not by ripping out an emitter, which is a two-repo change that also breaks a shared fixture. The implementation is its own session. And it is twenty facts, not twenty-one.
  • R-268 is closed — the leaked key is rotated, and the rotation is proved in both directions rather than assumed.

Fixed 2026-08-08 — four things the machine knew and did not say

A rebuilt machine can set up its own recovery again (R-221, agent 0.128.0 — proved on hardware); an unreadable disk is no longer drawn as a healthy empty one (R-259); a backup tick now answers about that app (R-258); our own alarm no longer points at a log that may not exist (R-265). Still true and not glossed: a failed disk reading still reaches us as "0 of 0 GB" — the quiet direction, it can only miss a true alarm, never raise a false one (R-266).

What we're working on

  • Widening the check so a fourth secret-in-a-page is caught by a machine. (R-255) · R-264 is now decided (above); building the readers is a session of its own.
  • Proving the hub really keeps the old sealed key when a machine re-seals. (R-198) · Still open, none urgent: (R-256, R-257, R-261…R-263, R-266)

Waiting on you

  • Nothing blocking. The rehearsal is finished and demo-hp is back in service: agent 0.128.0, controller 0.210.0, claimed, off-site backups unlocked and intact.
  • One decision worth taking before the tester comes: whether to fix the drive wall (R-280) now. It is the only finding that would stop his visit outright, and it is the difference between "his data comes back" and "his data comes back if someone types a path for him."
  • Two guards are still owed so the install cannot break the same way twice: refuse to vouch a version whose label does not resolve, and check that a published version and its label ship together. (R-273's tail.)

DooPlex infrastructure — separate from the product

Kept under its own heading rather than dropped: these are real asks that need you, but they concern the machine all this is built on, not what a customer receives. Mixing them in is why the page stopped being readable.

  • DooPlex's own backup keeps every copy inside the same box, and is silent when it fails. (R-232)
  • 193 old images exist only on this machine, ~27 GB — clutter, not space. (R-210)
  • The hub password needs rotating — a diagnostic printed it into a session log; nothing suggests anyone else saw it. (R-132)
  • After DooPlex next restarts, read /var/log/felhom-store-postboot-check.log — the second-SSD move has never survived a reboot; on PASS, 34 GB comes back. (R-209a)
  • Backup scripts on DooPlex are unversioned host state (R-231), and the instruction-file follow-ups each need a decision rather than an edit (R-229, R-230).
  • The Configuration page is fixed: 26 s → 0.14 s. It was never hashing anything — the hashes are already stored and simply read. It was making 42 calls one after another. Now they overlap, connections are kept, the answer is held for a minute, and the old artifacts are gone. Worst case is 5 s, once a minute at most. Your instinct to prune was right and my measurement said otherwise — trimming to ten of each halved the slow path. (R-267 — closed.)
  • The access token I printed into a log yesterday is rotated, and I checked it both ways: the old one is refused, the new one works, and the machine's own channel is back up. (R-268 — closed.)
  • Our build-check alarm has one gap left. A run that hangs is now cut off after five minutes and the mail says how long it took — but whether the alarm fires at all when the machinery kills a run outright is still unverified, and we have not claimed otherwise. (R-265)