STATUS back to one screen; close R-280 and R-294
gates / gates (push) Successful in 21s

211 lines -> one screen. Moves closed items out, corrects the tester paragraph,
states the floor situation as the operator's one-field call, and stops asking
him to decide something that shipped.
This commit is contained in:
2026-08-12 14:04:35 +02:00
parent 125aec1be2
commit 890a474ff2
2 changed files with 48 additions and 203 deletions
+46 -201
View File
@@ -1,211 +1,56 @@
# STATUS — what works, what's broken, what's next # STATUS — what works, what's broken, what's next
**Updated 2026-08-10 (afternoon).** **Updated 2026-08-12.**
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates > **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
> part of it in plain words, and **nothing may exist only here**. Not `CONTEXT.md`, which is technical > part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.**
> state written for Claude Code. **Items, not paragraphs. One screen.** If it does not fit, something > If it does not fit, it belongs in the register instead.
> belongs in the register instead.
>
> *Rebuilt from the register on 2026-08-07, from 258 lines. The old "what shipped recently" log is what
> the per-repo `CHANGELOG.md` files and the register are for, and is not restated here.*
## Controller 0.211.0 is baked, vouched and DELIVERED to fresh installs — the fleet is a separate switch
The tester-visit release shipped: the data drive can be **re-attached after a reinstall** (the restore
page's „két kattintás" was pointing at an empty picker — it was zero clicks); the orphan card **stops
promising** the set-aside off-site copies may be restorable, which the machine showing that card cannot
know; and the dashboard code is now called **„Beállító kód" everywhere** — „Visszaállító kód" is retired,
because it collided with the escrow „Helyreállítási kód" and that collision cost a real code.
Golden 0.211.0 is published and vouched (agent 0.128.0, min agent 0.127.0), verified by re-downloading
the served bytes and hashing them.
**One thing needs you.** The global update floor is **0.200.0**, and boxes auto-update to the *floor*,
never to the newest. So **fresh installs get these fixes and the existing machines do not** — demo-hp
and demo-felhom stay on 0.210.0 until the floor is raised. That is a one-field change on the same page,
and it is deliberately yours.
**Two things were dropped and are not forgotten:** our own uninstall still leaves `dnsmasq` holding
:53, so the next install refuses and blames the household's network (R-293 area, untouched); and the
hub's own emails still call the setup code by the retired name and send people to a page a rebuilt box
does not show (R-295, half done).
**The installer's stale-golden fix is written but NOT published** — an install could silently reuse an
old archive lying on the machine, including one too old to run the recovery screen. The fix is in
`main`, which publishes nothing; the tag is deliberately uncut until we have watched the failure happen
once on a drill machine (R-297).
## Both machines are home, unmuted and healthy — one thing still needs you
**Back online 2026-08-10 ~09:26 CEST**, both unblocked on the hub, both reporting **OK** on the
approved pair (agent 0.128.0, controller 0.210.0). No false alarm fired on power-up. `drill-r50` is
untouched and still blocked, as intended.
**demo-hp is in good shape.** Its off-site repository still opens with the machine's own key —
**18 snapshots, including yesterday's rehearsal files** — so the tier is credentialed and ready; its
first scheduled run since the rebuild is tonight at 04:15. Two apps it had before the rehearsal
(opengist, privatebin) were never reinstalled; only Calibre-Web was, as the walk needed.
**demo-felhom is protected again — and the recovery you authorised turned out to be impossible.** Before running it I checked, and the sealed package holds *the same key the machine already had* — a key that provably does not open its own backup store. Recovering it would have handed back something useless. The store was written under an older key whose sealed copy was **not retained** (the retention fix landed hours too late for it), so **those 1.2 GB are permanently unreadable by anyone, including us**.
So I took your stated fallback: the old store was **moved aside, not deleted** (`/home/felhom-repo.orphaned-20260810`), a fresh one was created under the current key, and a real backup ran — **succeeded in 10 seconds**, and I listed what is inside it rather than trusting the green tick: OpenGist's configuration, its manifest and its data volume. **The week without off-site protection is over.** *(R-278 closed.)*
**One thing that needs your judgement, not mine.** The card that offered this told the customer their set-aside backups *may be restorable later with their recovery code*. For these ones that is simply untrue, and it is said to precisely the people who have just lost their history. *(R-202 — now evidenced.)*
## Can anyone else lose their history the way demo-felhom did? No.
**One read of the hub's own records, no machine touched.** The hub holds backup keys for exactly
**three** machines. Both demo boxes lost their old key in the same four-hour window on 4 August,
before the retention fix was in force — that is the whole population of the problem, and it is
entirely ours. **The tester's machine has no record at all** — so it cannot be affected by *this*
defect, and that is the only reassuring thing about it: the reason it has no record is that its host
row was deleted on 15 July, and it has **no off-site copy, no key and no local backup either**. See
the `PETI` row in the register. Anything enrolled from now on is covered, because the fix has been in
force since 4 August.
I ran a control before trusting the query: it had to say *material present* for a machine known to
have it and *absent* for one known not to. It did both.
## Three green dots came back, nine stayed grey, and one rule finally fired
**The nine greys are the honest number.** For those, no document anywhere walks the claim, and saying
so is more useful than a dot nobody can defend.
- **Back to green**, each citing the document that walked it: the drive wizard (a live drive taken
through scan → format → mount → enrol), the on-box app backups (an overnight destructive campaign
across both machines), and the lost-recovery-code case — which we then proved the hard way this
morning.
- **One claim stayed grey for a new reason, and it is the interesting one.** The unattended
restore-proof *does* have a receipt from 28 July — but demo-hp's restore-test failed on 5 August and
the machine has since been wiped and rebuilt. It is a claim about something that keeps happening, so
an old observation cannot carry it. **This is the first time that rule has fired**; two nights ago it
fired zero times out of twelve.
## The prune mystery is solved, and the answer was written down all along
Who deleted the old versions: **you did, on 4 August evening, on your own rule** — 33 deletions, keep
set asserted first, every one a clean 204. **It was recorded inside the row about the Configuration
page being slow**, because pruning artifacts is what made that page fast. Two sessions failed to find
it. **You are no longer blocked** on establishing something that was already on file.
I also got a number wrong yesterday and it is corrected: I said the container packages held nineteen
versions and used that to argue against the prune. Counted properly — with pages — they hold 270 and
169, and the two that *were* pruned sit at exactly ten each.
## The sentence we should stop saying
When a machine's off-site history is set aside, the card tells the customer it *may be restorable
later with their recovery code*. **The machine showing that card cannot know whether it is true**
the fact lives on the hub and is not sent to the box. For anything set aside before 4 August it is
simply false. **The replacement wording is written and waiting**
(`documentation/design/SPEC-orphan-card-copy-2026-08-10.md`); it ships with the next controller
release so one image bake and one approval cover it, rather than costing you two of each.
## What works
A blank machine boots the Felhom disc, installs itself unattended, and is claimed by the customer, who
sets their own password. They install apps from a catalogue of fifty-three, share files over the home
network, and open apps from a launcher or a shared link. Backups run on their own to three places — the
machine's drive, a second drive, and an encrypted off-site copy.
**The backup promise is proved, and so is getting the data back yourself.** A machine has been
destroyed on purpose and its files came back byte for byte identical — four times now. On
**2026-08-07 the household's own journey passed for the first time**: someone with a browser and
their recovery code got everything back with **no command line inside the machine at any point**,
in 72 seconds. The two rough edges that walk found are also gone. *(R-201, R-252, R-253 — closed.)*
## Shipped 2026-08-09 — the guards, and the rehearsal that earned them
**An approval that cannot be installed is now refused** at the moment you press Save: the hub checks
the version's git label and that its file downloads, and refuses with a message naming the fix. A
second machine catches it a step earlier in the agent repo. Both were owed after every install in
existence failed for hours on 2026-08-09. **Hub v0.102.0 is live.**
**The reinstall rehearsal:** a demo machine was wiped and put back. All four test files returned
**byte for byte**, accented Hungarian filenames included, checked as raw bytes. But it only finished
because a terminal was available twice — the install died on a missing version label *(R-273, now
guarded)* and **a reinstalled machine still cannot re-attach its own data drive** *(R-280 — the one to
fix before the tester's visit)*. Detail: `documentation/audits/REHEARSAL-byo-reinstall-2026-08-09.md`.
## What's broken
- **Nothing else new is broken.** The *check* against a fourth secret-in-a-page covers 4 pages of 27, and
the cheap one covering all of them is blind to the shape that shipped. *(R-255)*
- **An already-paired box is still told to pair itself**, 25 minutes on. *(R-214, R-235)*
- **A backup that covered nothing still calls itself „Sikeres".** *(R-240)*
- **A machine waiting for its recovery code can stop backing up off-site without alarming us.** *(R-243)*
- **The card offering to reopen set-aside backups promises more than we can deliver.** *(R-202)*
- **Deleting a customer leaves rows behind** while reporting a clean teardown — no secrets, but it
accumulates. *(R-244)*
- **Putting restored files back where they belong is still manual.** *(R-213)*
## Three rulings, written down so they stop living in a conversation
- **The managed-update floor.** It is deliberately parked, and the trigger to raise it is **the first
machine that is not ours**; after that it moves with the publish train. Worth knowing alongside it:
the updater always aims at the floor, never at the newest, so a machine at or above the floor
updates to nothing. **Correction to the number that was going round: the floor is live at 0.200.0,
not 0.156.0** — checked twice today, on the hub page and in both machines' own logs.
- **R-264 is decided.** Build a reader for guest-network health, the staged-update pair, the
restore-test depth pair, and the two backup-integrity timestamps. Record a stated "no reader
wanted" for the repaired-recently flag, the tier-applied timestamp, the config fingerprint, the
per-stack object and the drive-migration marker. Decide the reporting-disabled flag on its own
merits — if that is a state we support, it must be visible or a staleness alarm will one day fire
on a machine that is fine. A "no" ends by changing the allowlist reason from *arguably owed* to
*deliberately not consumed* — not by ripping out an emitter, which is a two-repo change that also
breaks a shared fixture. **The implementation is its own session.** And **it is twenty facts, not
twenty-one**.
- **R-268 is closed** — the leaked key is rotated, and the rotation is proved in both directions
rather than assumed.
## Fixed 2026-08-08 — four things the machine knew and did not say
A rebuilt machine can set up its own recovery again *(R-221, agent 0.128.0 — proved on hardware)*; an
unreadable disk is no longer drawn as a healthy empty one *(R-259)*; a backup tick now answers about
*that* app *(R-258)*; our own alarm no longer points at a log that may not exist *(R-265)*. Still true
and not glossed: a failed disk reading still reaches us as "0 of 0 GB" — the quiet direction, it can
only miss a true alarm, never raise a false one *(R-266)*.
## What we're working on
- **Widening the check** so a fourth secret-in-a-page is caught by a machine. *(R-255)* · **R-264 is
now decided** (above); building the readers is a session of its own.
- **Proving the hub really keeps the old sealed key** when a machine re-seals. *(R-198)* · Still open,
none urgent: *(R-256, R-257, R-261…R-263, R-266)*
## Waiting on you ## Waiting on you
- **Nothing blocking.** The rehearsal is finished and demo-hp is back in service: agent 0.128.0, - **Raise the update floor to 0.212.0.** It is **0.200.0**, and boxes auto-update to the *floor*, never
controller 0.210.0, claimed, off-site backups unlocked and intact. to the newest — so everything below reaches **newly installed machines only**, and both demo boxes sit
- **One decision worth taking before the tester comes:** whether to fix the drive wall *(R-280)* now. where they are. One field, reversible. The hold that could block it does not apply (min agent 0.127.0
It is the only finding that would stop his visit outright, and it is the difference between "his is below the vouched agent 0.128.0). Recommended: do it and watch two disposable boxes move first.
data comes back" and "his data comes back if someone types a path for him." - **Vouch golden 0.212.0** (three fields: golden 0.212.0, agent 0.128.0, min agent 0.127.0).
- **Two guards are still owed** so the install cannot break the same way twice: refuse to vouch a - **R-296 / R-301 — two wording decisions** that need your call before anyone writes Hungarian at a
version whose label does not resolve, and check that a published version and its label ship customer. R-301 is the more interesting one: a banner still says the old promise, and it is probably
together. *(R-273's tail.)* true where it renders.
## DooPlex infrastructure — separate from the product ## What works
*Kept under its own heading rather than dropped: these are real asks that need you, but they concern Both demo machines are home, healthy and reporting on the approved pair. Off-site is credentialed on
the machine all this is built on, not what a customer receives. Mixing them in is why the page stopped `demo-hp` and its repository still opens with the machine's own key. `drill-r50` is blocked, as intended.
being readable.*
- **DooPlex's own backup keeps every copy inside the same box, and is silent when it fails.** *(R-232)* ## Shipped, delivered to fresh installs
- **193 old images exist only on this machine**, ~27 GB — clutter, not space. *(R-210)*
- **The hub password needs rotating** — a diagnostic printed it into a session log; nothing suggests - **The drive can be re-attached after a reinstall** (R-280). The restore page said *"this is two
anyone else saw it. *(R-132)* clicks"* over an empty list; it was zero clicks and needed an internal path no customer could produce.
- **After DooPlex next restarts**, read `/var/log/felhom-store-postboot-check.log` — the second-SSD - **The orphan card stops promising** that set-aside off-site copies can be reopened — twice over
move has never survived a reboot; on PASS, 34 GB comes back. *(R-209a)* (R-294, then **R-299**, which was the same claim in the plural, in the *always-visible* half, missed
- **Backup scripts on DooPlex are unversioned host state** *(R-231)*, and the instruction-file because the guard matched one inflection of a Hungarian verb).
follow-ups each need a decision rather than an edit *(R-229, R-230)*. - **One name per secret, box side** (R-295): the dashboard code is „Beállító kód" everywhere;
- **The Configuration page is fixed: 26 s → 0.14 s.** It was never hashing anything — the hashes „Visszaállító kód" is retired. It collided with the escrow „Helyreállítási kód" and cost a real code.
are already stored and simply read. It was making 42 calls one after another. Now they overlap,
connections are kept, the answer is held for a minute, and the old artifacts are gone. Worst case ## Broken, or knowingly incomplete
is 5 s, once a minute at most. **Your instinct to prune was right and my measurement said
otherwise** — trimming to ten of each halved the slow path. *(R-267 — closed.)* - **The tester's machine has no recovery route at all** — see the `PETI` row. Its host record was
- **The access token I printed into a log yesterday is rotated**, and I checked it both ways: the old deleted on 15 July; there is no key, no off-site copy and no local backup. **If that drive fails,
one is refused, the new one works, and the machine's own channel is back up. *(R-268 — closed.)* everything on it is lost.** First act of the visit: copy the ~3.6 GB off before anything is
- **Our build-check alarm has one gap left.** A run that hangs is now cut off after five minutes and reinstalled — it is currently the only copy in existence. Whether it stays parked is your call and is
the mail says how long it took — but **whether the alarm fires at all when the machinery kills a deliberately left open.
run outright is still unverified**, and we have not claimed otherwise. *(R-265)* - **Two installer fixes are written but NOT published** — pushing publishes nothing, and no tag is cut:
- **R-297** — an install could silently reuse an old base image lying on the machine, including one
too old to run the recovery screen.
- **R-300** — our own uninstall left `dnsmasq` holding `:53`, so our own next install refused and
blamed the household's network.
Both are unpublished for the same reason: **neither fault has been watched happening.** One session on
`drill-r50` covers both, and that is the right order.
- **The hub half of the naming is undone** (R-295 PARTIAL): the emails still use the retired name and
send people to a page a rebuilt machine does not show.
- **The storage page has its own separate reason for showing an empty list** (R-298), untouched.
## Working on next
The `drill-r50` session that unblocks both installer fixes; then the hub naming; then the
2026-08-09 batch (R-279 … R-292) which is still untriaged against everything since.
+2 -2
View File
@@ -508,7 +508,7 @@ applied.** The one that matters: Scenario A **fails against today's tree** with
| **R-277** | **Three hub surfaces jointly present a HEALTHY off-site tier as an absent one — and it produced a wrong operator statement during this run.** For demo-hp on 2026-08-09 the box was pushing off-site daily without a gap (18 restic snapshots, `last_status: ok`), yet: (a) the customer page's Backup panel read `Snapshots 0 / Repo Size 0 MB / Integrity Unknown` — it renders the **local disk tier**, while the healthy `offsite` object sits **in the same report** unrendered on that panel; (b) the Offsite page read `0.0 GB` — true, but a 162 KB repo rounds to nothing; (c) a stale `offsite_delivery_stuck` event from **2026-08-07 10:19** (not recurring) reads as current state. **Three independent surfaces agreeing on a wrong picture is how a working backup gets "fixed".** It did exactly that here: the rehearsal reported a fleet-wide off-site outage to the operator and had to retract it. **Note the true half:** demo-felhom IS genuinely stuck (`offsite.state=needs_credential`, no run has ever succeeded) → **R-278** | **READY (S) — NEW 2026-08-09** | — | Render the offsite object on the offsite row; show bytes not rounded GB; distinguish a live alarm from event history | CC | | **R-277** | **Three hub surfaces jointly present a HEALTHY off-site tier as an absent one — and it produced a wrong operator statement during this run.** For demo-hp on 2026-08-09 the box was pushing off-site daily without a gap (18 restic snapshots, `last_status: ok`), yet: (a) the customer page's Backup panel read `Snapshots 0 / Repo Size 0 MB / Integrity Unknown` — it renders the **local disk tier**, while the healthy `offsite` object sits **in the same report** unrendered on that panel; (b) the Offsite page read `0.0 GB` — true, but a 162 KB repo rounds to nothing; (c) a stale `offsite_delivery_stuck` event from **2026-08-07 10:19** (not recurring) reads as current state. **Three independent surfaces agreeing on a wrong picture is how a working backup gets "fixed".** It did exactly that here: the rehearsal reported a fleet-wide off-site outage to the operator and had to retract it. **Note the true half:** demo-felhom IS genuinely stuck (`offsite.state=needs_credential`, no run has ever succeeded) → **R-278** | **READY (S) — NEW 2026-08-09** | — | Render the offsite object on the offsite row; show bytes not rounded GB; distinguish a live alarm from event history | CC |
| **R-278** | **demo-felhom's off-site tier has never completed a run and has been stuck for six days.** `offsite.state=needs_credential` since the 2026-08-03 guest rebuild; the hub's own alarm reads *"enabled + escrowed but no run has EVER succeeded"*; the controller's `offsite-credential-retry` job runs every 5 minutes and completes in 0 s, doing nothing. R-193's fix (the recovery SCREEN, controller 0.200.0) is present on the box, so the remedy exists — it just needs the customer-present ceremony that nobody has run, which is R-243's shape (*"a machine waiting for its recovery code can stop backing up off-site without alarming us"*) landing on a real box. **Contrast that makes it a defect and not a chore:** demo-hp, same rebuild, same day, recovered and has 18 snapshots | **CLOSED 2026-08-10 — protection RESTORED, and the recovery it waited for could never have worked** | — | Either the self-heal reconciler owns this shape end-to-end, or the box must say plainly on the dashboard that it is unprotected pending the recovery code **THE REMEDY THIS ROW ASSUMED WAS IMPOSSIBLE, and that is the finding.** The operator authorised the recovery ceremony on 2026-08-10; it was **not run**, because three measurements taken first showed it could not help. The boxs local key hashes to `c60c8bc737a6b7c6…`; the hubs sealed escrow key hashes to **the same value**; and that key answers `Fatal: wrong password or no key found` against the repository. **Recovery would have returned a key the box already held and which was already proven not to open the store.** The repository was written under `48741892f0ef4d59…` (`host_escrow_superseded` id=4, superseded 2026-08-04 07:20:08) whose **`identity_blob` is NULL** — and the restic password lives ONLY in the identity bundle (`escrow/identity.go:39`, read by `escrow/recover.go:91` through `UnwrapIdentityBundle`), so it is unrecoverable by construction. Corroborated: the surviving K-escrow payload is **64 bytes**, a wrapped key, far too small to carry a bundle with a password. **The same shape the register already records for demo-hp** (*"its key sits in superseded row id 3 with identity_blob NULL … four hours before v0.93.0 fixed the retention"*). **What was done instead, on the operators stated fallback:** the orphan reset, through the customers own card — the old store **moved aside, never deleted**, to `/home/felhom-repo.orphaned-20260810` (1.2 GB); a fresh repository initialised under the current key; `offbox_repo_reset` audited hub-side at 08:06:31. **PROVEN rather than assumed:** `last_status: ok`, `last_success: 2026-08-10T08:07:33Z`, 10 s — and the snapshots CONTENTS listed, not just its count: `opengist/compose/{.felhom.yml,app.yaml,docker-compose.yml}`, `manifest.json`, and `volume-dumps/opengist_opengist_data.tar`. A week without off-site protection ends here | CC | | **R-278** | **demo-felhom's off-site tier has never completed a run and has been stuck for six days.** `offsite.state=needs_credential` since the 2026-08-03 guest rebuild; the hub's own alarm reads *"enabled + escrowed but no run has EVER succeeded"*; the controller's `offsite-credential-retry` job runs every 5 minutes and completes in 0 s, doing nothing. R-193's fix (the recovery SCREEN, controller 0.200.0) is present on the box, so the remedy exists — it just needs the customer-present ceremony that nobody has run, which is R-243's shape (*"a machine waiting for its recovery code can stop backing up off-site without alarming us"*) landing on a real box. **Contrast that makes it a defect and not a chore:** demo-hp, same rebuild, same day, recovered and has 18 snapshots | **CLOSED 2026-08-10 — protection RESTORED, and the recovery it waited for could never have worked** | — | Either the self-heal reconciler owns this shape end-to-end, or the box must say plainly on the dashboard that it is unprotected pending the recovery code **THE REMEDY THIS ROW ASSUMED WAS IMPOSSIBLE, and that is the finding.** The operator authorised the recovery ceremony on 2026-08-10; it was **not run**, because three measurements taken first showed it could not help. The boxs local key hashes to `c60c8bc737a6b7c6…`; the hubs sealed escrow key hashes to **the same value**; and that key answers `Fatal: wrong password or no key found` against the repository. **Recovery would have returned a key the box already held and which was already proven not to open the store.** The repository was written under `48741892f0ef4d59…` (`host_escrow_superseded` id=4, superseded 2026-08-04 07:20:08) whose **`identity_blob` is NULL** — and the restic password lives ONLY in the identity bundle (`escrow/identity.go:39`, read by `escrow/recover.go:91` through `UnwrapIdentityBundle`), so it is unrecoverable by construction. Corroborated: the surviving K-escrow payload is **64 bytes**, a wrapped key, far too small to carry a bundle with a password. **The same shape the register already records for demo-hp** (*"its key sits in superseded row id 3 with identity_blob NULL … four hours before v0.93.0 fixed the retention"*). **What was done instead, on the operators stated fallback:** the orphan reset, through the customers own card — the old store **moved aside, never deleted**, to `/home/felhom-repo.orphaned-20260810` (1.2 GB); a fresh repository initialised under the current key; `offbox_repo_reset` audited hub-side at 08:06:31. **PROVEN rather than assumed:** `last_status: ok`, `last_success: 2026-08-10T08:07:33Z`, 10 s — and the snapshots CONTENTS listed, not just its count: `opengist/compose/{.felhom.yml,app.yaml,docker-compose.yml}`, `manifest.json`, and `volume-dumps/opengist_opengist_data.tar`. A week without off-site protection ends here | CC |
| **R-279** | **There is no operator-triggerable off-site backup.** The only route to `POST /backup/offbox/run` is the customer's own dashboard session; `signed_jobs` carries opaque operator-SIGNED blobs and the hub holds no signing key. This cost the rehearsal a stop: preparing the run needed one off-site push and there was no operator path to it. Sibling of **R-177** (no operator-triggerable fill check) | **READY (XS) — NEW 2026-08-09** | — | Same shape as R-177; solve both together | CC | | **R-279** | **There is no operator-triggerable off-site backup.** The only route to `POST /backup/offbox/run` is the customer's own dashboard session; `signed_jobs` carries opaque operator-SIGNED blobs and the hub holds no signing key. This cost the rehearsal a stop: preparing the run needed one off-site push and there was no operator path to it. Sibling of **R-177** (no operator-triggerable fill check) | **READY (XS) — NEW 2026-08-09** | — | Same shape as R-177; solve both together | CC |
| **R-280** | **RANK 1 — after a reinstall the data drive cannot be re-attached through ANY dashboard route, and the restore page promises it is "two clicks".** Measured on the rebuilt demo-hp, 2026-08-09. The restore page diagnoses the situation perfectly and then sends the customer to an empty page: *„Előbb csatold vissza az adatmeghajtót. A mentéseid megvannak, és a meghajtók is megvannak — újratelepítés után viszont a gép még nem ismeri őket, ezért most nincs hová visszaállítani. **Ez két kattintás:** Tárhely → Meghajtók, »Meglévő meghajtó csatolása«."* **It is not two clicks; it is zero possible clicks.** `GET /api/disks/candidates``{"initialize":[],"attach":[]}`, so both wizards render an empty selector, and `Tárhely → Meghajtók` reads „Nincs regisztrált adattároló" with an empty unregistered list. **The agent is not at fault**`GET /api/disks` returns the NVMe in full (1.0 TB, SMART PASSED, `mount_path:/mnt/nvme-1tb`, `guest_attached:false`), so the channel and enumeration work. **ROOT CAUSE:** `handleDiskCandidates` builds both lists from `ListCandidateDisks`, the UNCLAIMED-disk scan; demo-hp's NVMe is deliberately BOTH the user-data drive and the `felhom-backup` target (`operations/nodes.md`), so it is claimed and never offered. That filter is **correct for `initialize`** (never offer to format a disk in use — `/storage/init` even says so: *„Rendszer- és biztonsági-mentés meghajtók itt nem jelennek meg — azok védettek"*) and **over-broad for `attach`**, which is non-destructive by definition and whose own page says *„A meghajtón lévő adatok nem törlődnek — a csatolás csak elérhetővé teszi azokat."* **It cascades:** no store → Calibre-Web's install page degrades to *„Nincs regisztrált adattároló — adja meg kézzel az útvonalat"* and demands a hand-typed `E-könyvtár útvonal`; no app → the restore rows read „Nincs telepítve". **THE ESCAPE HATCH WORKS AND NO CUSTOMER COULD FIND IT:** `POST /settings/storage/add` with `storage_path=/mnt/sys_drive` succeeded first try (*„Adattároló sikeresen hozzáadva"*) — and `/mnt/sys_drive` is an internal path, the very one registered before the wipe. Once registered, everything unblocked and the deploy form became a proper picker (*„Tárhely (sys_drive) — 64.2 GB szabad"*). **This is R-220's successor:** R-220 was closed as "drives unenrollable after a rebuild — fixed"; enumeration is fixed, OFFERING is not | **READY (M) — NEW 2026-08-09** | — | Populate `attach` from mounted-but-unregistered filesystems rather than from the unclaimed-DISK scan; and never print "two clicks" without asserting the destination is non-empty | CC | | **R-280** | **RANK 1 — after a reinstall the data drive cannot be re-attached through ANY dashboard route, and the restore page promises it is "two clicks".** Measured on the rebuilt demo-hp, 2026-08-09. The restore page diagnoses the situation perfectly and then sends the customer to an empty page: *„Előbb csatold vissza az adatmeghajtót. A mentéseid megvannak, és a meghajtók is megvannak — újratelepítés után viszont a gép még nem ismeri őket, ezért most nincs hová visszaállítani. **Ez két kattintás:** Tárhely → Meghajtók, »Meglévő meghajtó csatolása«."* **It is not two clicks; it is zero possible clicks.** `GET /api/disks/candidates``{"initialize":[],"attach":[]}`, so both wizards render an empty selector, and `Tárhely → Meghajtók` reads „Nincs regisztrált adattároló" with an empty unregistered list. **The agent is not at fault**`GET /api/disks` returns the NVMe in full (1.0 TB, SMART PASSED, `mount_path:/mnt/nvme-1tb`, `guest_attached:false`), so the channel and enumeration work. **ROOT CAUSE:** `handleDiskCandidates` builds both lists from `ListCandidateDisks`, the UNCLAIMED-disk scan; demo-hp's NVMe is deliberately BOTH the user-data drive and the `felhom-backup` target (`operations/nodes.md`), so it is claimed and never offered. That filter is **correct for `initialize`** (never offer to format a disk in use — `/storage/init` even says so: *„Rendszer- és biztonsági-mentés meghajtók itt nem jelennek meg — azok védettek"*) and **over-broad for `attach`**, which is non-destructive by definition and whose own page says *„A meghajtón lévő adatok nem törlődnek — a csatolás csak elérhetővé teszi azokat."* **It cascades:** no store → Calibre-Web's install page degrades to *„Nincs regisztrált adattároló — adja meg kézzel az útvonalat"* and demands a hand-typed `E-könyvtár útvonal`; no app → the restore rows read „Nincs telepítve". **THE ESCAPE HATCH WORKS AND NO CUSTOMER COULD FIND IT:** `POST /settings/storage/add` with `storage_path=/mnt/sys_drive` succeeded first try (*„Adattároló sikeresen hozzáadva"*) — and `/mnt/sys_drive` is an internal path, the very one registered before the wipe. Once registered, everything unblocked and the deploy form became a proper picker (*„Tárhely (sys_drive) — 64.2 GB szabad"*). **This is R-220's successor:** R-220 was closed as "drives unenrollable after a rebuild — fixed"; enumeration is fixed, OFFERING is not | **CLOSED — controller v0.211.0, delivered via golden 0.211.0 (vouched 2026-08-10)** | — | Populate `attach` from mounted-but-unregistered filesystems rather than from the unclaimed-DISK scan; and never print "two clicks" without asserting the destination is non-empty | CC |
| **R-281** | ~~**The hub said NOTHING through an entire reinstall — and the tripwire for a sealed-backup unseal did not fire on a real unseal.**~~ **WITHDRAWN 2026-08-09 — THE FINDING WAS AN ARTEFACT OF MY OWN MEASUREMENT, AND IT WAS WRONG IN BOTH DIRECTIONS.** The operator's mailbox settled it: the hub fired **twenty events** on 2026-08-09, and `escrow_blob_served` **DID** fire — 10:19:41 UTC / **12:19 CEST**, eight minutes before the verified restore. **Cause of the false reading, ESTABLISHED (not guessed):** the P7 query copied `/data/hub.db` **without `hub.db-wal`**. The hub runs SQLite in WAL mode (R-172), so every write since the last checkpoint was invisible. **The signature is an exact match:** P7 reported *"2 events all day, newest `db_dump_completed` 00:30:07"*, and the number of rows on 08-09 at or before 00:30:07 is **exactly 2**. **The two obvious alternatives were TESTED AND REFUTED**, not waved away: a **timezone offset** — all nine mailbox stamps equal the hub's UTC + 2 h exactly (`escrow_blob_served` 10:19→12:19, `host_down` 09:28→11:28, and seven more), so the window was right; and a **wrong customer key or wrong store** — the same table and key return the correct rows now. A live re-run cannot reproduce the fault because the WAL has since been checkpointed; the case rests on the command text plus the 2-of-2 count signature, and that is stated rather than dressed up as a reproduction. **This is a trap this project has already documented**`operations/nodes.md` says copying `hub.db` alone is *"valid but stale … the worst failure shape"* — and I had avoided it correctly earlier in the same session before hitting it. **Split out: → R-285** (the real, opposite defect) and **→ R-286** (the measurement lesson) | **WITHDRAWN 2026-08-09** | — | Superseded by R-285/R-286 | CC | | **R-281** | ~~**The hub said NOTHING through an entire reinstall — and the tripwire for a sealed-backup unseal did not fire on a real unseal.**~~ **WITHDRAWN 2026-08-09 — THE FINDING WAS AN ARTEFACT OF MY OWN MEASUREMENT, AND IT WAS WRONG IN BOTH DIRECTIONS.** The operator's mailbox settled it: the hub fired **twenty events** on 2026-08-09, and `escrow_blob_served` **DID** fire — 10:19:41 UTC / **12:19 CEST**, eight minutes before the verified restore. **Cause of the false reading, ESTABLISHED (not guessed):** the P7 query copied `/data/hub.db` **without `hub.db-wal`**. The hub runs SQLite in WAL mode (R-172), so every write since the last checkpoint was invisible. **The signature is an exact match:** P7 reported *"2 events all day, newest `db_dump_completed` 00:30:07"*, and the number of rows on 08-09 at or before 00:30:07 is **exactly 2**. **The two obvious alternatives were TESTED AND REFUTED**, not waved away: a **timezone offset** — all nine mailbox stamps equal the hub's UTC + 2 h exactly (`escrow_blob_served` 10:19→12:19, `host_down` 09:28→11:28, and seven more), so the window was right; and a **wrong customer key or wrong store** — the same table and key return the correct rows now. A live re-run cannot reproduce the fault because the WAL has since been checkpointed; the case rests on the command text plus the 2-of-2 count signature, and that is stated rather than dressed up as a reproduction. **This is a trap this project has already documented**`operations/nodes.md` says copying `hub.db` alone is *"valid but stale … the worst failure shape"* — and I had avoided it correctly earlier in the same session before hitting it. **Split out: → R-285** (the real, opposite defect) and **→ R-286** (the measurement lesson) | **WITHDRAWN 2026-08-09** | — | Superseded by R-285/R-286 | CC |
| **R-282** | **One secret, three different Hungarian names, and the email sends the customer to a page their box is not showing.** Sending it from the hub is „**Visszaállító** kód küldése"; the email that arrives is subject „Jelszó-**visszaállítási** kód", body „**Visszaállító** kód: …", and it instructs *„Add meg a vezérlőpult »**Elfelejtett jelszó**« oldalán"*; the page the box actually serves is „A szerver **beállítása**" asking for a „**Beállító** kód". **A rebuilt box shows a SETUP page and the hub can only send a RESET mail** (because hub-side the customer is still `claimed_at 2026-07-21`), so the instruction names a route that does not exist on screen. **It does work if you ignore the instructions** — the reset code was accepted on the setup page (302 + session), so this is naming, not function. **It cost this session real time and one wasted code:** the operator supplied a 3-word Hungarian code believing it was the recovery code, because the hub calls the claim code „Visszaállító kód" and the ESCROW code is also „Visszaállító kód" — the only reliable discriminator is length (claim = 3 Hungarian words; recovery = **10** EFF-list words, and the recovery screen does say „(tíz szó)") | **READY (S) — NEW 2026-08-09** | — | Pick one name per secret and use it on all three surfaces; make the mail's page reference match what a rebuilt box actually shows | CC | | **R-282** | **One secret, three different Hungarian names, and the email sends the customer to a page their box is not showing.** Sending it from the hub is „**Visszaállító** kód küldése"; the email that arrives is subject „Jelszó-**visszaállítási** kód", body „**Visszaállító** kód: …", and it instructs *„Add meg a vezérlőpult »**Elfelejtett jelszó**« oldalán"*; the page the box actually serves is „A szerver **beállítása**" asking for a „**Beállító** kód". **A rebuilt box shows a SETUP page and the hub can only send a RESET mail** (because hub-side the customer is still `claimed_at 2026-07-21`), so the instruction names a route that does not exist on screen. **It does work if you ignore the instructions** — the reset code was accepted on the setup page (302 + session), so this is naming, not function. **It cost this session real time and one wasted code:** the operator supplied a 3-word Hungarian code believing it was the recovery code, because the hub calls the claim code „Visszaállító kód" and the ESCROW code is also „Visszaállító kód" — the only reliable discriminator is length (claim = 3 Hungarian words; recovery = **10** EFF-list words, and the recovery screen does say „(tíz szó)") | **READY (S) — NEW 2026-08-09** | — | Pick one name per secret and use it on all three surfaces; make the mail's page reference match what a rebuilt box actually shows | CC |
| **R-283** | **After a rebuild the hub says "Claimed 18d ago" while the box serves its first-run setup page.** `customer_claims` for demo-hp still read `claimed_at 2026-07-21 16:29:25`, `generation 2`, `issued_at 2026-08-03` while the freshly provisioned guest — whose `settings.json` is new — correctly showed „A szerver beállítása". The two sides never reconcile: the hub's claim state survives a guest rebuild and the box's does not. Consequences: the operator's screen says the box is claimed when it is not, a resend produces a RESET code instead of a SETUP code (→ **R-282**), and any previously issued code fails with *„Hibás vagy lejárt kód"* — a message that is technically true and tells the customer nothing about the real cause, namely their own reinstall. Mirror image of **R-214/R-235** (an already-paired box still told to pair itself) | **READY (S) — NEW 2026-08-09** | — | Let a report from a box carrying no claim state clear the hub's, or show both sides on the operator page | CC | | **R-283** | **After a rebuild the hub says "Claimed 18d ago" while the box serves its first-run setup page.** `customer_claims` for demo-hp still read `claimed_at 2026-07-21 16:29:25`, `generation 2`, `issued_at 2026-08-03` while the freshly provisioned guest — whose `settings.json` is new — correctly showed „A szerver beállítása". The two sides never reconcile: the hub's claim state survives a guest rebuild and the box's does not. Consequences: the operator's screen says the box is claimed when it is not, a resend produces a RESET code instead of a SETUP code (→ **R-282**), and any previously issued code fails with *„Hibás vagy lejárt kód"* — a message that is technically true and tells the customer nothing about the real cause, namely their own reinstall. Mirror image of **R-214/R-235** (an already-paired box still told to pair itself) | **READY (S) — NEW 2026-08-09** | — | Let a report from a box carrying no claim state clear the hub's, or show both sides on the operator page | CC |
@@ -522,7 +522,7 @@ applied.** The one that matters: Scenario A **fails against today's tree** with
| **R-291** | **CI's installability assertion is now BOUNDED by a retention number, and the narrowing is recorded here so it can be widened deliberately rather than discovered.** `check-published-versions.py` demanded that **every** `v<semver>` tag still be downloadable while the registry demonstrably does not retain every version — two sensible rules that cannot both hold, which is why CI went red at a commit whose own run had been green the day before, and would have gone red again at the next publish. **The fix couples them:** `felhom-agent/scripts/retention-policy.json` is THE number (`generic_versions_kept: 10`) and the check reads it. **WHAT CI NO LONGER COVERS, stated plainly: a released version older than the retention window is no longer asserted downloadable.** Its git TAG and its config tree are still asserted — only the binary's presence is dropped — and the check **prints the dropped versions on every run**, so the narrowing cannot go quiet. Controls run: widened to 11 the evicted version re-enters and convicts (exit 1); the policy file removed gives INCONCLUSIVE (exit 2), never silently unbounded. **The number is an OBSERVED state, not a located ruling** (R-287) and the file says so. **The better bound, recorded rather than built:** the hub's vouched `min_agent` floor — nothing can install an agent below it, so a sub-floor version being un-downloadable costs nothing real; it needs the gate to read the hub, which is network it does not have today | **READY (S) — NEW 2026-08-09** | R-287 | ~~Widen or replace the number when the deleter is established~~**CONDITION RELEASED 2026-08-10: the deleter IS established (R-287), so the operator is no longer blocked on establishing what was already written down.** The number can now be confirmed or replaced on its merits. The better bound remains the vouched `min_agent` floor | CC | | **R-291** | **CI's installability assertion is now BOUNDED by a retention number, and the narrowing is recorded here so it can be widened deliberately rather than discovered.** `check-published-versions.py` demanded that **every** `v<semver>` tag still be downloadable while the registry demonstrably does not retain every version — two sensible rules that cannot both hold, which is why CI went red at a commit whose own run had been green the day before, and would have gone red again at the next publish. **The fix couples them:** `felhom-agent/scripts/retention-policy.json` is THE number (`generic_versions_kept: 10`) and the check reads it. **WHAT CI NO LONGER COVERS, stated plainly: a released version older than the retention window is no longer asserted downloadable.** Its git TAG and its config tree are still asserted — only the binary's presence is dropped — and the check **prints the dropped versions on every run**, so the narrowing cannot go quiet. Controls run: widened to 11 the evicted version re-enters and convicts (exit 1); the policy file removed gives INCONCLUSIVE (exit 2), never silently unbounded. **The number is an OBSERVED state, not a located ruling** (R-287) and the file says so. **The better bound, recorded rather than built:** the hub's vouched `min_agent` floor — nothing can install an agent below it, so a sub-floor version being un-downloadable costs nothing real; it needs the gate to read the hub, which is network it does not have today | **READY (S) — NEW 2026-08-09** | R-287 | ~~Widen or replace the number when the deleter is established~~**CONDITION RELEASED 2026-08-10: the deleter IS established (R-287), so the operator is no longer blocked on establishing what was already written down.** The number can now be confirmed or replaced on its merits. The better bound remains the vouched `min_agent` floor | CC |
| **R-292** | **The artifact-save flash conflates three different facts, and a failing test found it rather than a reading.** `artifact_sha_invalid` reads *"the Gitea sha lookup failed (version missing / Gitea unreachable) or the manually-entered sha is invalid"* — three causes, one message, and the operator acts differently on each. It surfaced because scenario E of the new installability gate kept reporting `artifact_sha_invalid` where it expected `artifact_unverifiable`: `resolveArtifactSHA` ran first and swallowed the distinction. **Worked around in v0.102.0 by ORDERING** — the installability probes now run before the sha resolution, so an unreachable registry is reported as unreachable — **but the underlying message is untouched and still conflates on its own paths** | **READY (XS) — NEW 2026-08-09** | — | Split it into "version not found", "registry unreachable" and "invalid sha" | CC | | **R-292** | **The artifact-save flash conflates three different facts, and a failing test found it rather than a reading.** `artifact_sha_invalid` reads *"the Gitea sha lookup failed (version missing / Gitea unreachable) or the manually-entered sha is invalid"* — three causes, one message, and the operator acts differently on each. It surfaced because scenario E of the new installability gate kept reporting `artifact_sha_invalid` where it expected `artifact_unverifiable`: `resolveArtifactSHA` ran first and swallowed the distinction. **Worked around in v0.102.0 by ORDERING** — the installability probes now run before the sha resolution, so an unreachable registry is reported as unreachable — **but the underlying message is untouched and still conflates on its own paths** | **READY (XS) — NEW 2026-08-09** | — | Split it into "version not found", "registry unreachable" and "invalid sha" | CC |
| **R-293** | **CENSUS, 2026-08-10 — no machine that is not ours can be in the state that cost demo-felhom its history, and here is the whole population.** Read-only against the hub store, with a control run first (the query returned *present (572 bytes)* for a host known to have material and *absent (NULL)* for one known not to — both cases from the two demo boxes, so the instrument was shown to distinguish the states before it was trusted). **The hub knows three hosts.** `demo-felhom-8363b5` and `demo-hp-bb76ea` each hold one superseded escrow with **`identity_blob` ABSENT**, superseded **07:20:08** and **07:15:36** on 2026-08-04 — both **before** the retention fix was in force, pinned at **11:11:37Z** from the hub's own first post-fix escrow row (a date-only comparison mislabels these as "after" and was corrected). `drill-r50-0a4f9a` has no supersession. **`peti-felhom` — the tester's machine — has NO host row and NO escrow at all**, and neither does `david`; the orphan check found no escrow row pointing at an unknown host. **So the answer is: no, not today, and not tomorrow either** — any future enrolment escrows under the fixed code. **What this does NOT claim:** that a retained blob has ever been *unwrapped* on a superseded row. Retention is proven; the recovery FROM a superseded row is still unexercised | **CLOSED-INFORMATIONAL 2026-08-10** | — | The machine was not contacted; only the hub's records were read | CC | | **R-293** | **CENSUS, 2026-08-10 — no machine that is not ours can be in the state that cost demo-felhom its history, and here is the whole population.** Read-only against the hub store, with a control run first (the query returned *present (572 bytes)* for a host known to have material and *absent (NULL)* for one known not to — both cases from the two demo boxes, so the instrument was shown to distinguish the states before it was trusted). **The hub knows three hosts.** `demo-felhom-8363b5` and `demo-hp-bb76ea` each hold one superseded escrow with **`identity_blob` ABSENT**, superseded **07:20:08** and **07:15:36** on 2026-08-04 — both **before** the retention fix was in force, pinned at **11:11:37Z** from the hub's own first post-fix escrow row (a date-only comparison mislabels these as "after" and was corrected). `drill-r50-0a4f9a` has no supersession. **`peti-felhom` — the tester's machine — has NO host row and NO escrow at all**, and neither does `david`; the orphan check found no escrow row pointing at an unknown host. **So the answer is: no, not today, and not tomorrow either** — any future enrolment escrows under the fixed code. **What this does NOT claim:** that a retained blob has ever been *unwrapped* on a superseded row. Retention is proven; the recovery FROM a superseded row is still unexercised | **CLOSED-INFORMATIONAL 2026-08-10** | — | The machine was not contacted; only the hub's records were read | CC |
| **R-294** | **The orphan card promises restorability that the box rendering it cannot evaluate — specified, not implemented.** The card says the set-aside copies *"a hozzá tartozó helyreállítási kóddal később visszaállítható lehet"* (`controller/internal/web/templates/backups_remote.html:101`). **The discriminator lives in the hub** (`host_escrow_superseded.identity_blob`, `hub/internal/store/store.go:393`); **the box caches only `HubEscrowIdentityPresent`** (`controller/internal/settings/settings.go:71`), which describes the CURRENT escrow, not a superseded one; and **no field on the report or ACK wire carries superseded-blob retention**. So the renderer cannot tell which case the customer is in — **a conditional promise the system cannot evaluate is the same defect as an unconditional false one.** Replacement Hungarian copy, the surfaces at `file:line`, and the render tests that should pin it: **`documentation/design/SPEC-orphan-card-copy-2026-08-10.md`**. **Deliberately NOT implemented:** it lands in the controller, and a controller release is undelivered until a golden carries it (R-242) — one bake, one approval. **It should ship with the next controller change so one bake covers both**, and the spec says so. Parent: **R-202** | **READY (S) — NEW 2026-08-10** | R-202 | Implement with the next controller release, not on its own | CC | | **R-294** | **The orphan card promises restorability that the box rendering it cannot evaluate — specified, not implemented.** The card says the set-aside copies *"a hozzá tartozó helyreállítási kóddal később visszaállítható lehet"* (`controller/internal/web/templates/backups_remote.html:101`). **The discriminator lives in the hub** (`host_escrow_superseded.identity_blob`, `hub/internal/store/store.go:393`); **the box caches only `HubEscrowIdentityPresent`** (`controller/internal/settings/settings.go:71`), which describes the CURRENT escrow, not a superseded one; and **no field on the report or ACK wire carries superseded-blob retention**. So the renderer cannot tell which case the customer is in — **a conditional promise the system cannot evaluate is the same defect as an unconditional false one.** Replacement Hungarian copy, the surfaces at `file:line`, and the render tests that should pin it: **`documentation/design/SPEC-orphan-card-copy-2026-08-10.md`**. **Deliberately NOT implemented:** it lands in the controller, and a controller release is undelivered until a golden carries it (R-242) — one bake, one approval. **It should ship with the next controller change so one bake covers both**, and the spec says so. Parent: **R-202** | **CLOSED — controller v0.211.0; see R-299 for the sentence it missed** | R-202 | Implement with the next controller release, not on its own | CC |
**Explicitly still open, untouched by this session:** R-246, R-255, R-256, R-257, R-261, R-262, **Explicitly still open, untouched by this session:** R-246, R-255, R-256, R-257, R-261, R-262,
R-263, **R-264** (the twenty-one undecided facts — a design session of its own), R-240, R-243, R-263, **R-264** (the twenty-one undecided facts — a design session of its own), R-240, R-243,