Files
felhom.eu/STATUS.md
T
2026-08-10 14:29:54 +02:00

209 lines
15 KiB
Markdown

# STATUS — what works, what's broken, what's next
**Updated 2026-08-10 (afternoon).**
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
> part of it in plain words, and **nothing may exist only here**. Not `CONTEXT.md`, which is technical
> state written for Claude Code. **Items, not paragraphs. One screen.** If it does not fit, something
> belongs in the register instead.
>
> *Rebuilt from the register on 2026-08-07, from 258 lines. The old "what shipped recently" log is what
> the per-repo `CHANGELOG.md` files and the register are for, and is not restated here.*
## Controller 0.211.0 is baked, vouched and DELIVERED to fresh installs — the fleet is a separate switch
The tester-visit release shipped: the data drive can be **re-attached after a reinstall** (the restore
page's „két kattintás" was pointing at an empty picker — it was zero clicks); the orphan card **stops
promising** the set-aside off-site copies may be restorable, which the machine showing that card cannot
know; and the dashboard code is now called **„Beállító kód" everywhere** — „Visszaállító kód" is retired,
because it collided with the escrow „Helyreállítási kód" and that collision cost a real code.
Golden 0.211.0 is published and vouched (agent 0.128.0, min agent 0.127.0), verified by re-downloading
the served bytes and hashing them.
**One thing needs you.** The global update floor is **0.200.0**, and boxes auto-update to the *floor*,
never to the newest. So **fresh installs get these fixes and the existing machines do not** — demo-hp
and demo-felhom stay on 0.210.0 until the floor is raised. That is a one-field change on the same page,
and it is deliberately yours.
**Two things were dropped and are not forgotten:** our own uninstall still leaves `dnsmasq` holding
:53, so the next install refuses and blames the household's network (R-293 area, untouched); and the
hub's own emails still call the setup code by the retired name and send people to a page a rebuilt box
does not show (R-295, half done).
**The installer's stale-golden fix is written but NOT published** — an install could silently reuse an
old archive lying on the machine, including one too old to run the recovery screen. The fix is in
`main`, which publishes nothing; the tag is deliberately uncut until we have watched the failure happen
once on a drill machine (R-297).
## Both machines are home, unmuted and healthy — one thing still needs you
**Back online 2026-08-10 ~09:26 CEST**, both unblocked on the hub, both reporting **OK** on the
approved pair (agent 0.128.0, controller 0.210.0). No false alarm fired on power-up. `drill-r50` is
untouched and still blocked, as intended.
**demo-hp is in good shape.** Its off-site repository still opens with the machine's own key —
**18 snapshots, including yesterday's rehearsal files** — so the tier is credentialed and ready; its
first scheduled run since the rebuild is tonight at 04:15. Two apps it had before the rehearsal
(opengist, privatebin) were never reinstalled; only Calibre-Web was, as the walk needed.
**demo-felhom is protected again — and the recovery you authorised turned out to be impossible.** Before running it I checked, and the sealed package holds *the same key the machine already had* — a key that provably does not open its own backup store. Recovering it would have handed back something useless. The store was written under an older key whose sealed copy was **not retained** (the retention fix landed hours too late for it), so **those 1.2 GB are permanently unreadable by anyone, including us**.
So I took your stated fallback: the old store was **moved aside, not deleted** (`/home/felhom-repo.orphaned-20260810`), a fresh one was created under the current key, and a real backup ran — **succeeded in 10 seconds**, and I listed what is inside it rather than trusting the green tick: OpenGist's configuration, its manifest and its data volume. **The week without off-site protection is over.** *(R-278 closed.)*
**One thing that needs your judgement, not mine.** The card that offered this told the customer their set-aside backups *may be restorable later with their recovery code*. For these ones that is simply untrue, and it is said to precisely the people who have just lost their history. *(R-202 — now evidenced.)*
## Can anyone else lose their history the way demo-felhom did? No.
**One read of the hub's own records, no machine touched.** The hub holds backup keys for exactly
**three** machines. Both demo boxes lost their old key in the same four-hour window on 4 August,
before the retention fix was in force — that is the whole population of the problem, and it is
entirely ours. **The tester's machine has no record at all**, so it cannot be affected; and anything
enrolled from now on is covered, because the fix has been in force since 4 August.
I ran a control before trusting the query: it had to say *material present* for a machine known to
have it and *absent* for one known not to. It did both.
## Three green dots came back, nine stayed grey, and one rule finally fired
**The nine greys are the honest number.** For those, no document anywhere walks the claim, and saying
so is more useful than a dot nobody can defend.
- **Back to green**, each citing the document that walked it: the drive wizard (a live drive taken
through scan → format → mount → enrol), the on-box app backups (an overnight destructive campaign
across both machines), and the lost-recovery-code case — which we then proved the hard way this
morning.
- **One claim stayed grey for a new reason, and it is the interesting one.** The unattended
restore-proof *does* have a receipt from 28 July — but demo-hp's restore-test failed on 5 August and
the machine has since been wiped and rebuilt. It is a claim about something that keeps happening, so
an old observation cannot carry it. **This is the first time that rule has fired**; two nights ago it
fired zero times out of twelve.
## The prune mystery is solved, and the answer was written down all along
Who deleted the old versions: **you did, on 4 August evening, on your own rule** — 33 deletions, keep
set asserted first, every one a clean 204. **It was recorded inside the row about the Configuration
page being slow**, because pruning artifacts is what made that page fast. Two sessions failed to find
it. **You are no longer blocked** on establishing something that was already on file.
I also got a number wrong yesterday and it is corrected: I said the container packages held nineteen
versions and used that to argue against the prune. Counted properly — with pages — they hold 270 and
169, and the two that *were* pruned sit at exactly ten each.
## The sentence we should stop saying
When a machine's off-site history is set aside, the card tells the customer it *may be restorable
later with their recovery code*. **The machine showing that card cannot know whether it is true**
the fact lives on the hub and is not sent to the box. For anything set aside before 4 August it is
simply false. **The replacement wording is written and waiting**
(`documentation/design/SPEC-orphan-card-copy-2026-08-10.md`); it ships with the next controller
release so one image bake and one approval cover it, rather than costing you two of each.
## What works
A blank machine boots the Felhom disc, installs itself unattended, and is claimed by the customer, who
sets their own password. They install apps from a catalogue of fifty-three, share files over the home
network, and open apps from a launcher or a shared link. Backups run on their own to three places — the
machine's drive, a second drive, and an encrypted off-site copy.
**The backup promise is proved, and so is getting the data back yourself.** A machine has been
destroyed on purpose and its files came back byte for byte identical — four times now. On
**2026-08-07 the household's own journey passed for the first time**: someone with a browser and
their recovery code got everything back with **no command line inside the machine at any point**,
in 72 seconds. The two rough edges that walk found are also gone. *(R-201, R-252, R-253 — closed.)*
## Shipped 2026-08-09 — the guards, and the rehearsal that earned them
**An approval that cannot be installed is now refused** at the moment you press Save: the hub checks
the version's git label and that its file downloads, and refuses with a message naming the fix. A
second machine catches it a step earlier in the agent repo. Both were owed after every install in
existence failed for hours on 2026-08-09. **Hub v0.102.0 is live.**
**The reinstall rehearsal:** a demo machine was wiped and put back. All four test files returned
**byte for byte**, accented Hungarian filenames included, checked as raw bytes. But it only finished
because a terminal was available twice — the install died on a missing version label *(R-273, now
guarded)* and **a reinstalled machine still cannot re-attach its own data drive** *(R-280 — the one to
fix before the tester's visit)*. Detail: `documentation/audits/REHEARSAL-byo-reinstall-2026-08-09.md`.
## What's broken
- **Nothing else new is broken.** The *check* against a fourth secret-in-a-page covers 4 pages of 27, and
the cheap one covering all of them is blind to the shape that shipped. *(R-255)*
- **An already-paired box is still told to pair itself**, 25 minutes on. *(R-214, R-235)*
- **A backup that covered nothing still calls itself „Sikeres".** *(R-240)*
- **A machine waiting for its recovery code can stop backing up off-site without alarming us.** *(R-243)*
- **The card offering to reopen set-aside backups promises more than we can deliver.** *(R-202)*
- **Deleting a customer leaves rows behind** while reporting a clean teardown — no secrets, but it
accumulates. *(R-244)*
- **Putting restored files back where they belong is still manual.** *(R-213)*
## Three rulings, written down so they stop living in a conversation
- **The managed-update floor.** It is deliberately parked, and the trigger to raise it is **the first
machine that is not ours**; after that it moves with the publish train. Worth knowing alongside it:
the updater always aims at the floor, never at the newest, so a machine at or above the floor
updates to nothing. **Correction to the number that was going round: the floor is live at 0.200.0,
not 0.156.0** — checked twice today, on the hub page and in both machines' own logs.
- **R-264 is decided.** Build a reader for guest-network health, the staged-update pair, the
restore-test depth pair, and the two backup-integrity timestamps. Record a stated "no reader
wanted" for the repaired-recently flag, the tier-applied timestamp, the config fingerprint, the
per-stack object and the drive-migration marker. Decide the reporting-disabled flag on its own
merits — if that is a state we support, it must be visible or a staleness alarm will one day fire
on a machine that is fine. A "no" ends by changing the allowlist reason from *arguably owed* to
*deliberately not consumed* — not by ripping out an emitter, which is a two-repo change that also
breaks a shared fixture. **The implementation is its own session.** And **it is twenty facts, not
twenty-one**.
- **R-268 is closed** — the leaked key is rotated, and the rotation is proved in both directions
rather than assumed.
## Fixed 2026-08-08 — four things the machine knew and did not say
A rebuilt machine can set up its own recovery again *(R-221, agent 0.128.0 — proved on hardware)*; an
unreadable disk is no longer drawn as a healthy empty one *(R-259)*; a backup tick now answers about
*that* app *(R-258)*; our own alarm no longer points at a log that may not exist *(R-265)*. Still true
and not glossed: a failed disk reading still reaches us as "0 of 0 GB" — the quiet direction, it can
only miss a true alarm, never raise a false one *(R-266)*.
## What we're working on
- **Widening the check** so a fourth secret-in-a-page is caught by a machine. *(R-255)* · **R-264 is
now decided** (above); building the readers is a session of its own.
- **Proving the hub really keeps the old sealed key** when a machine re-seals. *(R-198)* · Still open,
none urgent: *(R-256, R-257, R-261…R-263, R-266)*
## Waiting on you
- **Nothing blocking.** The rehearsal is finished and demo-hp is back in service: agent 0.128.0,
controller 0.210.0, claimed, off-site backups unlocked and intact.
- **One decision worth taking before the tester comes:** whether to fix the drive wall *(R-280)* now.
It is the only finding that would stop his visit outright, and it is the difference between "his
data comes back" and "his data comes back if someone types a path for him."
- **Two guards are still owed** so the install cannot break the same way twice: refuse to vouch a
version whose label does not resolve, and check that a published version and its label ship
together. *(R-273's tail.)*
## DooPlex infrastructure — separate from the product
*Kept under its own heading rather than dropped: these are real asks that need you, but they concern
the machine all this is built on, not what a customer receives. Mixing them in is why the page stopped
being readable.*
- **DooPlex's own backup keeps every copy inside the same box, and is silent when it fails.** *(R-232)*
- **193 old images exist only on this machine**, ~27 GB — clutter, not space. *(R-210)*
- **The hub password needs rotating** — a diagnostic printed it into a session log; nothing suggests
anyone else saw it. *(R-132)*
- **After DooPlex next restarts**, read `/var/log/felhom-store-postboot-check.log` — the second-SSD
move has never survived a reboot; on PASS, 34 GB comes back. *(R-209a)*
- **Backup scripts on DooPlex are unversioned host state** *(R-231)*, and the instruction-file
follow-ups each need a decision rather than an edit *(R-229, R-230)*.
- **The Configuration page is fixed: 26 s → 0.14 s.** It was never hashing anything — the hashes
are already stored and simply read. It was making 42 calls one after another. Now they overlap,
connections are kept, the answer is held for a minute, and the old artifacts are gone. Worst case
is 5 s, once a minute at most. **Your instinct to prune was right and my measurement said
otherwise** — trimming to ten of each halved the slow path. *(R-267 — closed.)*
- **The access token I printed into a log yesterday is rotated**, and I checked it both ways: the old
one is refused, the new one works, and the machine's own channel is back up. *(R-268 — closed.)*
- **Our build-check alarm has one gap left.** A run that hangs is now cut off after five minutes and
the mail says how long it took — but **whether the alarm fires at all when the machinery kills a
run outright is still unverified**, and we have not claimed otherwise. *(R-265)*