209 lines
15 KiB
Markdown
209 lines
15 KiB
Markdown
# STATUS — what works, what's broken, what's next
|
|
|
|
**Updated 2026-08-10 (afternoon).**
|
|
|
|
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
|
|
> part of it in plain words, and **nothing may exist only here**. Not `CONTEXT.md`, which is technical
|
|
> state written for Claude Code. **Items, not paragraphs. One screen.** If it does not fit, something
|
|
> belongs in the register instead.
|
|
>
|
|
> *Rebuilt from the register on 2026-08-07, from 258 lines. The old "what shipped recently" log is what
|
|
> the per-repo `CHANGELOG.md` files and the register are for, and is not restated here.*
|
|
|
|
## Controller 0.211.0 is baked, vouched and DELIVERED to fresh installs — the fleet is a separate switch
|
|
|
|
The tester-visit release shipped: the data drive can be **re-attached after a reinstall** (the restore
|
|
page's „két kattintás" was pointing at an empty picker — it was zero clicks); the orphan card **stops
|
|
promising** the set-aside off-site copies may be restorable, which the machine showing that card cannot
|
|
know; and the dashboard code is now called **„Beállító kód" everywhere** — „Visszaállító kód" is retired,
|
|
because it collided with the escrow „Helyreállítási kód" and that collision cost a real code.
|
|
|
|
Golden 0.211.0 is published and vouched (agent 0.128.0, min agent 0.127.0), verified by re-downloading
|
|
the served bytes and hashing them.
|
|
|
|
**One thing needs you.** The global update floor is **0.200.0**, and boxes auto-update to the *floor*,
|
|
never to the newest. So **fresh installs get these fixes and the existing machines do not** — demo-hp
|
|
and demo-felhom stay on 0.210.0 until the floor is raised. That is a one-field change on the same page,
|
|
and it is deliberately yours.
|
|
|
|
**Two things were dropped and are not forgotten:** our own uninstall still leaves `dnsmasq` holding
|
|
:53, so the next install refuses and blames the household's network (R-293 area, untouched); and the
|
|
hub's own emails still call the setup code by the retired name and send people to a page a rebuilt box
|
|
does not show (R-295, half done).
|
|
|
|
**The installer's stale-golden fix is written but NOT published** — an install could silently reuse an
|
|
old archive lying on the machine, including one too old to run the recovery screen. The fix is in
|
|
`main`, which publishes nothing; the tag is deliberately uncut until we have watched the failure happen
|
|
once on a drill machine (R-297).
|
|
|
|
## Both machines are home, unmuted and healthy — one thing still needs you
|
|
|
|
**Back online 2026-08-10 ~09:26 CEST**, both unblocked on the hub, both reporting **OK** on the
|
|
approved pair (agent 0.128.0, controller 0.210.0). No false alarm fired on power-up. `drill-r50` is
|
|
untouched and still blocked, as intended.
|
|
|
|
**demo-hp is in good shape.** Its off-site repository still opens with the machine's own key —
|
|
**18 snapshots, including yesterday's rehearsal files** — so the tier is credentialed and ready; its
|
|
first scheduled run since the rebuild is tonight at 04:15. Two apps it had before the rehearsal
|
|
(opengist, privatebin) were never reinstalled; only Calibre-Web was, as the walk needed.
|
|
|
|
**demo-felhom is protected again — and the recovery you authorised turned out to be impossible.** Before running it I checked, and the sealed package holds *the same key the machine already had* — a key that provably does not open its own backup store. Recovering it would have handed back something useless. The store was written under an older key whose sealed copy was **not retained** (the retention fix landed hours too late for it), so **those 1.2 GB are permanently unreadable by anyone, including us**.
|
|
|
|
So I took your stated fallback: the old store was **moved aside, not deleted** (`/home/felhom-repo.orphaned-20260810`), a fresh one was created under the current key, and a real backup ran — **succeeded in 10 seconds**, and I listed what is inside it rather than trusting the green tick: OpenGist's configuration, its manifest and its data volume. **The week without off-site protection is over.** *(R-278 closed.)*
|
|
|
|
**One thing that needs your judgement, not mine.** The card that offered this told the customer their set-aside backups *may be restorable later with their recovery code*. For these ones that is simply untrue, and it is said to precisely the people who have just lost their history. *(R-202 — now evidenced.)*
|
|
|
|
## Can anyone else lose their history the way demo-felhom did? No.
|
|
|
|
**One read of the hub's own records, no machine touched.** The hub holds backup keys for exactly
|
|
**three** machines. Both demo boxes lost their old key in the same four-hour window on 4 August,
|
|
before the retention fix was in force — that is the whole population of the problem, and it is
|
|
entirely ours. **The tester's machine has no record at all**, so it cannot be affected; and anything
|
|
enrolled from now on is covered, because the fix has been in force since 4 August.
|
|
|
|
I ran a control before trusting the query: it had to say *material present* for a machine known to
|
|
have it and *absent* for one known not to. It did both.
|
|
|
|
## Three green dots came back, nine stayed grey, and one rule finally fired
|
|
|
|
**The nine greys are the honest number.** For those, no document anywhere walks the claim, and saying
|
|
so is more useful than a dot nobody can defend.
|
|
|
|
- **Back to green**, each citing the document that walked it: the drive wizard (a live drive taken
|
|
through scan → format → mount → enrol), the on-box app backups (an overnight destructive campaign
|
|
across both machines), and the lost-recovery-code case — which we then proved the hard way this
|
|
morning.
|
|
- **One claim stayed grey for a new reason, and it is the interesting one.** The unattended
|
|
restore-proof *does* have a receipt from 28 July — but demo-hp's restore-test failed on 5 August and
|
|
the machine has since been wiped and rebuilt. It is a claim about something that keeps happening, so
|
|
an old observation cannot carry it. **This is the first time that rule has fired**; two nights ago it
|
|
fired zero times out of twelve.
|
|
|
|
## The prune mystery is solved, and the answer was written down all along
|
|
|
|
Who deleted the old versions: **you did, on 4 August evening, on your own rule** — 33 deletions, keep
|
|
set asserted first, every one a clean 204. **It was recorded inside the row about the Configuration
|
|
page being slow**, because pruning artifacts is what made that page fast. Two sessions failed to find
|
|
it. **You are no longer blocked** on establishing something that was already on file.
|
|
|
|
I also got a number wrong yesterday and it is corrected: I said the container packages held nineteen
|
|
versions and used that to argue against the prune. Counted properly — with pages — they hold 270 and
|
|
169, and the two that *were* pruned sit at exactly ten each.
|
|
|
|
## The sentence we should stop saying
|
|
|
|
When a machine's off-site history is set aside, the card tells the customer it *may be restorable
|
|
later with their recovery code*. **The machine showing that card cannot know whether it is true** —
|
|
the fact lives on the hub and is not sent to the box. For anything set aside before 4 August it is
|
|
simply false. **The replacement wording is written and waiting**
|
|
(`documentation/design/SPEC-orphan-card-copy-2026-08-10.md`); it ships with the next controller
|
|
release so one image bake and one approval cover it, rather than costing you two of each.
|
|
|
|
## What works
|
|
|
|
A blank machine boots the Felhom disc, installs itself unattended, and is claimed by the customer, who
|
|
sets their own password. They install apps from a catalogue of fifty-three, share files over the home
|
|
network, and open apps from a launcher or a shared link. Backups run on their own to three places — the
|
|
machine's drive, a second drive, and an encrypted off-site copy.
|
|
|
|
**The backup promise is proved, and so is getting the data back yourself.** A machine has been
|
|
destroyed on purpose and its files came back byte for byte identical — four times now. On
|
|
**2026-08-07 the household's own journey passed for the first time**: someone with a browser and
|
|
their recovery code got everything back with **no command line inside the machine at any point**,
|
|
in 72 seconds. The two rough edges that walk found are also gone. *(R-201, R-252, R-253 — closed.)*
|
|
|
|
## Shipped 2026-08-09 — the guards, and the rehearsal that earned them
|
|
|
|
**An approval that cannot be installed is now refused** at the moment you press Save: the hub checks
|
|
the version's git label and that its file downloads, and refuses with a message naming the fix. A
|
|
second machine catches it a step earlier in the agent repo. Both were owed after every install in
|
|
existence failed for hours on 2026-08-09. **Hub v0.102.0 is live.**
|
|
|
|
**The reinstall rehearsal:** a demo machine was wiped and put back. All four test files returned
|
|
**byte for byte**, accented Hungarian filenames included, checked as raw bytes. But it only finished
|
|
because a terminal was available twice — the install died on a missing version label *(R-273, now
|
|
guarded)* and **a reinstalled machine still cannot re-attach its own data drive** *(R-280 — the one to
|
|
fix before the tester's visit)*. Detail: `documentation/audits/REHEARSAL-byo-reinstall-2026-08-09.md`.
|
|
|
|
## What's broken
|
|
|
|
- **Nothing else new is broken.** The *check* against a fourth secret-in-a-page covers 4 pages of 27, and
|
|
the cheap one covering all of them is blind to the shape that shipped. *(R-255)*
|
|
- **An already-paired box is still told to pair itself**, 25 minutes on. *(R-214, R-235)*
|
|
- **A backup that covered nothing still calls itself „Sikeres".** *(R-240)*
|
|
- **A machine waiting for its recovery code can stop backing up off-site without alarming us.** *(R-243)*
|
|
- **The card offering to reopen set-aside backups promises more than we can deliver.** *(R-202)*
|
|
- **Deleting a customer leaves rows behind** while reporting a clean teardown — no secrets, but it
|
|
accumulates. *(R-244)*
|
|
- **Putting restored files back where they belong is still manual.** *(R-213)*
|
|
|
|
## Three rulings, written down so they stop living in a conversation
|
|
|
|
- **The managed-update floor.** It is deliberately parked, and the trigger to raise it is **the first
|
|
machine that is not ours**; after that it moves with the publish train. Worth knowing alongside it:
|
|
the updater always aims at the floor, never at the newest, so a machine at or above the floor
|
|
updates to nothing. **Correction to the number that was going round: the floor is live at 0.200.0,
|
|
not 0.156.0** — checked twice today, on the hub page and in both machines' own logs.
|
|
- **R-264 is decided.** Build a reader for guest-network health, the staged-update pair, the
|
|
restore-test depth pair, and the two backup-integrity timestamps. Record a stated "no reader
|
|
wanted" for the repaired-recently flag, the tier-applied timestamp, the config fingerprint, the
|
|
per-stack object and the drive-migration marker. Decide the reporting-disabled flag on its own
|
|
merits — if that is a state we support, it must be visible or a staleness alarm will one day fire
|
|
on a machine that is fine. A "no" ends by changing the allowlist reason from *arguably owed* to
|
|
*deliberately not consumed* — not by ripping out an emitter, which is a two-repo change that also
|
|
breaks a shared fixture. **The implementation is its own session.** And **it is twenty facts, not
|
|
twenty-one**.
|
|
- **R-268 is closed** — the leaked key is rotated, and the rotation is proved in both directions
|
|
rather than assumed.
|
|
|
|
## Fixed 2026-08-08 — four things the machine knew and did not say
|
|
|
|
A rebuilt machine can set up its own recovery again *(R-221, agent 0.128.0 — proved on hardware)*; an
|
|
unreadable disk is no longer drawn as a healthy empty one *(R-259)*; a backup tick now answers about
|
|
*that* app *(R-258)*; our own alarm no longer points at a log that may not exist *(R-265)*. Still true
|
|
and not glossed: a failed disk reading still reaches us as "0 of 0 GB" — the quiet direction, it can
|
|
only miss a true alarm, never raise a false one *(R-266)*.
|
|
|
|
## What we're working on
|
|
|
|
- **Widening the check** so a fourth secret-in-a-page is caught by a machine. *(R-255)* · **R-264 is
|
|
now decided** (above); building the readers is a session of its own.
|
|
- **Proving the hub really keeps the old sealed key** when a machine re-seals. *(R-198)* · Still open,
|
|
none urgent: *(R-256, R-257, R-261…R-263, R-266)*
|
|
|
|
## Waiting on you
|
|
|
|
- **Nothing blocking.** The rehearsal is finished and demo-hp is back in service: agent 0.128.0,
|
|
controller 0.210.0, claimed, off-site backups unlocked and intact.
|
|
- **One decision worth taking before the tester comes:** whether to fix the drive wall *(R-280)* now.
|
|
It is the only finding that would stop his visit outright, and it is the difference between "his
|
|
data comes back" and "his data comes back if someone types a path for him."
|
|
- **Two guards are still owed** so the install cannot break the same way twice: refuse to vouch a
|
|
version whose label does not resolve, and check that a published version and its label ship
|
|
together. *(R-273's tail.)*
|
|
|
|
## DooPlex infrastructure — separate from the product
|
|
|
|
*Kept under its own heading rather than dropped: these are real asks that need you, but they concern
|
|
the machine all this is built on, not what a customer receives. Mixing them in is why the page stopped
|
|
being readable.*
|
|
|
|
- **DooPlex's own backup keeps every copy inside the same box, and is silent when it fails.** *(R-232)*
|
|
- **193 old images exist only on this machine**, ~27 GB — clutter, not space. *(R-210)*
|
|
- **The hub password needs rotating** — a diagnostic printed it into a session log; nothing suggests
|
|
anyone else saw it. *(R-132)*
|
|
- **After DooPlex next restarts**, read `/var/log/felhom-store-postboot-check.log` — the second-SSD
|
|
move has never survived a reboot; on PASS, 34 GB comes back. *(R-209a)*
|
|
- **Backup scripts on DooPlex are unversioned host state** *(R-231)*, and the instruction-file
|
|
follow-ups each need a decision rather than an edit *(R-229, R-230)*.
|
|
- **The Configuration page is fixed: 26 s → 0.14 s.** It was never hashing anything — the hashes
|
|
are already stored and simply read. It was making 42 calls one after another. Now they overlap,
|
|
connections are kept, the answer is held for a minute, and the old artifacts are gone. Worst case
|
|
is 5 s, once a minute at most. **Your instinct to prune was right and my measurement said
|
|
otherwise** — trimming to ten of each halved the slow path. *(R-267 — closed.)*
|
|
- **The access token I printed into a log yesterday is rotated**, and I checked it both ways: the old
|
|
one is refused, the new one works, and the machine's own channel is back up. *(R-268 — closed.)*
|
|
- **Our build-check alarm has one gap left.** A run that hangs is now cut off after five minutes and
|
|
the mail says how long it took — but **whether the alarm fires at all when the machinery kills a
|
|
run outright is still unverified**, and we have not claimed otherwise. *(R-265)*
|