From 890a474ff2a4cbdbc620bc1dd3ae63e6600d8024 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Wed, 12 Aug 2026 14:04:35 +0200 Subject: [PATCH] STATUS back to one screen; close R-280 and R-294 211 lines -> one screen. Moves closed items out, corrects the tester paragraph, states the floor situation as the operator's one-field call, and stops asking him to decide something that shipped. --- STATUS.md | 247 ++++++---------------------- documentation/backlog/OPEN-ITEMS.md | 4 +- 2 files changed, 48 insertions(+), 203 deletions(-) diff --git a/STATUS.md b/STATUS.md index a4238f8..ac55ac8 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,211 +1,56 @@ # STATUS — what works, what's broken, what's next -**Updated 2026-08-10 (afternoon).** +**Updated 2026-08-12.** > **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates -> part of it in plain words, and **nothing may exist only here**. Not `CONTEXT.md`, which is technical -> state written for Claude Code. **Items, not paragraphs. One screen.** If it does not fit, something -> belongs in the register instead. -> -> *Rebuilt from the register on 2026-08-07, from 258 lines. The old "what shipped recently" log is what -> the per-repo `CHANGELOG.md` files and the register are for, and is not restated here.* - -## Controller 0.211.0 is baked, vouched and DELIVERED to fresh installs — the fleet is a separate switch - -The tester-visit release shipped: the data drive can be **re-attached after a reinstall** (the restore -page's „két kattintás" was pointing at an empty picker — it was zero clicks); the orphan card **stops -promising** the set-aside off-site copies may be restorable, which the machine showing that card cannot -know; and the dashboard code is now called **„Beállító kód" everywhere** — „Visszaállító kód" is retired, -because it collided with the escrow „Helyreállítási kód" and that collision cost a real code. - -Golden 0.211.0 is published and vouched (agent 0.128.0, min agent 0.127.0), verified by re-downloading -the served bytes and hashing them. - -**One thing needs you.** The global update floor is **0.200.0**, and boxes auto-update to the *floor*, -never to the newest. So **fresh installs get these fixes and the existing machines do not** — demo-hp -and demo-felhom stay on 0.210.0 until the floor is raised. That is a one-field change on the same page, -and it is deliberately yours. - -**Two things were dropped and are not forgotten:** our own uninstall still leaves `dnsmasq` holding -:53, so the next install refuses and blames the household's network (R-293 area, untouched); and the -hub's own emails still call the setup code by the retired name and send people to a page a rebuilt box -does not show (R-295, half done). - -**The installer's stale-golden fix is written but NOT published** — an install could silently reuse an -old archive lying on the machine, including one too old to run the recovery screen. The fix is in -`main`, which publishes nothing; the tag is deliberately uncut until we have watched the failure happen -once on a drill machine (R-297). - -## Both machines are home, unmuted and healthy — one thing still needs you - -**Back online 2026-08-10 ~09:26 CEST**, both unblocked on the hub, both reporting **OK** on the -approved pair (agent 0.128.0, controller 0.210.0). No false alarm fired on power-up. `drill-r50` is -untouched and still blocked, as intended. - -**demo-hp is in good shape.** Its off-site repository still opens with the machine's own key — -**18 snapshots, including yesterday's rehearsal files** — so the tier is credentialed and ready; its -first scheduled run since the rebuild is tonight at 04:15. Two apps it had before the rehearsal -(opengist, privatebin) were never reinstalled; only Calibre-Web was, as the walk needed. - -**demo-felhom is protected again — and the recovery you authorised turned out to be impossible.** Before running it I checked, and the sealed package holds *the same key the machine already had* — a key that provably does not open its own backup store. Recovering it would have handed back something useless. The store was written under an older key whose sealed copy was **not retained** (the retention fix landed hours too late for it), so **those 1.2 GB are permanently unreadable by anyone, including us**. - -So I took your stated fallback: the old store was **moved aside, not deleted** (`/home/felhom-repo.orphaned-20260810`), a fresh one was created under the current key, and a real backup ran — **succeeded in 10 seconds**, and I listed what is inside it rather than trusting the green tick: OpenGist's configuration, its manifest and its data volume. **The week without off-site protection is over.** *(R-278 closed.)* - -**One thing that needs your judgement, not mine.** The card that offered this told the customer their set-aside backups *may be restorable later with their recovery code*. For these ones that is simply untrue, and it is said to precisely the people who have just lost their history. *(R-202 — now evidenced.)* - -## Can anyone else lose their history the way demo-felhom did? No. - -**One read of the hub's own records, no machine touched.** The hub holds backup keys for exactly -**three** machines. Both demo boxes lost their old key in the same four-hour window on 4 August, -before the retention fix was in force — that is the whole population of the problem, and it is -entirely ours. **The tester's machine has no record at all** — so it cannot be affected by *this* -defect, and that is the only reassuring thing about it: the reason it has no record is that its host -row was deleted on 15 July, and it has **no off-site copy, no key and no local backup either**. See -the `PETI` row in the register. Anything enrolled from now on is covered, because the fix has been in -force since 4 August. - -I ran a control before trusting the query: it had to say *material present* for a machine known to -have it and *absent* for one known not to. It did both. - -## Three green dots came back, nine stayed grey, and one rule finally fired - -**The nine greys are the honest number.** For those, no document anywhere walks the claim, and saying -so is more useful than a dot nobody can defend. - -- **Back to green**, each citing the document that walked it: the drive wizard (a live drive taken - through scan → format → mount → enrol), the on-box app backups (an overnight destructive campaign - across both machines), and the lost-recovery-code case — which we then proved the hard way this - morning. -- **One claim stayed grey for a new reason, and it is the interesting one.** The unattended - restore-proof *does* have a receipt from 28 July — but demo-hp's restore-test failed on 5 August and - the machine has since been wiped and rebuilt. It is a claim about something that keeps happening, so - an old observation cannot carry it. **This is the first time that rule has fired**; two nights ago it - fired zero times out of twelve. - -## The prune mystery is solved, and the answer was written down all along - -Who deleted the old versions: **you did, on 4 August evening, on your own rule** — 33 deletions, keep -set asserted first, every one a clean 204. **It was recorded inside the row about the Configuration -page being slow**, because pruning artifacts is what made that page fast. Two sessions failed to find -it. **You are no longer blocked** on establishing something that was already on file. - -I also got a number wrong yesterday and it is corrected: I said the container packages held nineteen -versions and used that to argue against the prune. Counted properly — with pages — they hold 270 and -169, and the two that *were* pruned sit at exactly ten each. - -## The sentence we should stop saying - -When a machine's off-site history is set aside, the card tells the customer it *may be restorable -later with their recovery code*. **The machine showing that card cannot know whether it is true** — -the fact lives on the hub and is not sent to the box. For anything set aside before 4 August it is -simply false. **The replacement wording is written and waiting** -(`documentation/design/SPEC-orphan-card-copy-2026-08-10.md`); it ships with the next controller -release so one image bake and one approval cover it, rather than costing you two of each. - -## What works - -A blank machine boots the Felhom disc, installs itself unattended, and is claimed by the customer, who -sets their own password. They install apps from a catalogue of fifty-three, share files over the home -network, and open apps from a launcher or a shared link. Backups run on their own to three places — the -machine's drive, a second drive, and an encrypted off-site copy. - -**The backup promise is proved, and so is getting the data back yourself.** A machine has been -destroyed on purpose and its files came back byte for byte identical — four times now. On -**2026-08-07 the household's own journey passed for the first time**: someone with a browser and -their recovery code got everything back with **no command line inside the machine at any point**, -in 72 seconds. The two rough edges that walk found are also gone. *(R-201, R-252, R-253 — closed.)* - -## Shipped 2026-08-09 — the guards, and the rehearsal that earned them - -**An approval that cannot be installed is now refused** at the moment you press Save: the hub checks -the version's git label and that its file downloads, and refuses with a message naming the fix. A -second machine catches it a step earlier in the agent repo. Both were owed after every install in -existence failed for hours on 2026-08-09. **Hub v0.102.0 is live.** - -**The reinstall rehearsal:** a demo machine was wiped and put back. All four test files returned -**byte for byte**, accented Hungarian filenames included, checked as raw bytes. But it only finished -because a terminal was available twice — the install died on a missing version label *(R-273, now -guarded)* and **a reinstalled machine still cannot re-attach its own data drive** *(R-280 — the one to -fix before the tester's visit)*. Detail: `documentation/audits/REHEARSAL-byo-reinstall-2026-08-09.md`. - -## What's broken - -- **Nothing else new is broken.** The *check* against a fourth secret-in-a-page covers 4 pages of 27, and - the cheap one covering all of them is blind to the shape that shipped. *(R-255)* -- **An already-paired box is still told to pair itself**, 25 minutes on. *(R-214, R-235)* -- **A backup that covered nothing still calls itself „Sikeres".** *(R-240)* -- **A machine waiting for its recovery code can stop backing up off-site without alarming us.** *(R-243)* -- **The card offering to reopen set-aside backups promises more than we can deliver.** *(R-202)* -- **Deleting a customer leaves rows behind** while reporting a clean teardown — no secrets, but it - accumulates. *(R-244)* -- **Putting restored files back where they belong is still manual.** *(R-213)* - -## Three rulings, written down so they stop living in a conversation - -- **The managed-update floor.** It is deliberately parked, and the trigger to raise it is **the first - machine that is not ours**; after that it moves with the publish train. Worth knowing alongside it: - the updater always aims at the floor, never at the newest, so a machine at or above the floor - updates to nothing. **Correction to the number that was going round: the floor is live at 0.200.0, - not 0.156.0** — checked twice today, on the hub page and in both machines' own logs. -- **R-264 is decided.** Build a reader for guest-network health, the staged-update pair, the - restore-test depth pair, and the two backup-integrity timestamps. Record a stated "no reader - wanted" for the repaired-recently flag, the tier-applied timestamp, the config fingerprint, the - per-stack object and the drive-migration marker. Decide the reporting-disabled flag on its own - merits — if that is a state we support, it must be visible or a staleness alarm will one day fire - on a machine that is fine. A "no" ends by changing the allowlist reason from *arguably owed* to - *deliberately not consumed* — not by ripping out an emitter, which is a two-repo change that also - breaks a shared fixture. **The implementation is its own session.** And **it is twenty facts, not - twenty-one**. -- **R-268 is closed** — the leaked key is rotated, and the rotation is proved in both directions - rather than assumed. - -## Fixed 2026-08-08 — four things the machine knew and did not say - -A rebuilt machine can set up its own recovery again *(R-221, agent 0.128.0 — proved on hardware)*; an -unreadable disk is no longer drawn as a healthy empty one *(R-259)*; a backup tick now answers about -*that* app *(R-258)*; our own alarm no longer points at a log that may not exist *(R-265)*. Still true -and not glossed: a failed disk reading still reaches us as "0 of 0 GB" — the quiet direction, it can -only miss a true alarm, never raise a false one *(R-266)*. - -## What we're working on - -- **Widening the check** so a fourth secret-in-a-page is caught by a machine. *(R-255)* · **R-264 is - now decided** (above); building the readers is a session of its own. -- **Proving the hub really keeps the old sealed key** when a machine re-seals. *(R-198)* · Still open, - none urgent: *(R-256, R-257, R-261…R-263, R-266)* +> part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.** +> If it does not fit, it belongs in the register instead. ## Waiting on you -- **Nothing blocking.** The rehearsal is finished and demo-hp is back in service: agent 0.128.0, - controller 0.210.0, claimed, off-site backups unlocked and intact. -- **One decision worth taking before the tester comes:** whether to fix the drive wall *(R-280)* now. - It is the only finding that would stop his visit outright, and it is the difference between "his - data comes back" and "his data comes back if someone types a path for him." -- **Two guards are still owed** so the install cannot break the same way twice: refuse to vouch a - version whose label does not resolve, and check that a published version and its label ship - together. *(R-273's tail.)* +- **Raise the update floor to 0.212.0.** It is **0.200.0**, and boxes auto-update to the *floor*, never + to the newest — so everything below reaches **newly installed machines only**, and both demo boxes sit + where they are. One field, reversible. The hold that could block it does not apply (min agent 0.127.0 + is below the vouched agent 0.128.0). Recommended: do it and watch two disposable boxes move first. +- **Vouch golden 0.212.0** (three fields: golden 0.212.0, agent 0.128.0, min agent 0.127.0). +- **R-296 / R-301 — two wording decisions** that need your call before anyone writes Hungarian at a + customer. R-301 is the more interesting one: a banner still says the old promise, and it is probably + true where it renders. -## DooPlex infrastructure — separate from the product +## What works -*Kept under its own heading rather than dropped: these are real asks that need you, but they concern -the machine all this is built on, not what a customer receives. Mixing them in is why the page stopped -being readable.* +Both demo machines are home, healthy and reporting on the approved pair. Off-site is credentialed on +`demo-hp` and its repository still opens with the machine's own key. `drill-r50` is blocked, as intended. -- **DooPlex's own backup keeps every copy inside the same box, and is silent when it fails.** *(R-232)* -- **193 old images exist only on this machine**, ~27 GB — clutter, not space. *(R-210)* -- **The hub password needs rotating** — a diagnostic printed it into a session log; nothing suggests - anyone else saw it. *(R-132)* -- **After DooPlex next restarts**, read `/var/log/felhom-store-postboot-check.log` — the second-SSD - move has never survived a reboot; on PASS, 34 GB comes back. *(R-209a)* -- **Backup scripts on DooPlex are unversioned host state** *(R-231)*, and the instruction-file - follow-ups each need a decision rather than an edit *(R-229, R-230)*. -- **The Configuration page is fixed: 26 s → 0.14 s.** It was never hashing anything — the hashes - are already stored and simply read. It was making 42 calls one after another. Now they overlap, - connections are kept, the answer is held for a minute, and the old artifacts are gone. Worst case - is 5 s, once a minute at most. **Your instinct to prune was right and my measurement said - otherwise** — trimming to ten of each halved the slow path. *(R-267 — closed.)* -- **The access token I printed into a log yesterday is rotated**, and I checked it both ways: the old - one is refused, the new one works, and the machine's own channel is back up. *(R-268 — closed.)* -- **Our build-check alarm has one gap left.** A run that hangs is now cut off after five minutes and - the mail says how long it took — but **whether the alarm fires at all when the machinery kills a - run outright is still unverified**, and we have not claimed otherwise. *(R-265)* +## Shipped, delivered to fresh installs + +- **The drive can be re-attached after a reinstall** (R-280). The restore page said *"this is two + clicks"* over an empty list; it was zero clicks and needed an internal path no customer could produce. +- **The orphan card stops promising** that set-aside off-site copies can be reopened — twice over + (R-294, then **R-299**, which was the same claim in the plural, in the *always-visible* half, missed + because the guard matched one inflection of a Hungarian verb). +- **One name per secret, box side** (R-295): the dashboard code is „Beállító kód" everywhere; + „Visszaállító kód" is retired. It collided with the escrow „Helyreállítási kód" and cost a real code. + +## Broken, or knowingly incomplete + +- **The tester's machine has no recovery route at all** — see the `PETI` row. Its host record was + deleted on 15 July; there is no key, no off-site copy and no local backup. **If that drive fails, + everything on it is lost.** First act of the visit: copy the ~3.6 GB off before anything is + reinstalled — it is currently the only copy in existence. Whether it stays parked is your call and is + deliberately left open. +- **Two installer fixes are written but NOT published** — pushing publishes nothing, and no tag is cut: + - **R-297** — an install could silently reuse an old base image lying on the machine, including one + too old to run the recovery screen. + - **R-300** — our own uninstall left `dnsmasq` holding `:53`, so our own next install refused and + blamed the household's network. + Both are unpublished for the same reason: **neither fault has been watched happening.** One session on + `drill-r50` covers both, and that is the right order. +- **The hub half of the naming is undone** (R-295 PARTIAL): the emails still use the retired name and + send people to a page a rebuilt machine does not show. +- **The storage page has its own separate reason for showing an empty list** (R-298), untouched. + +## Working on next + +The `drill-r50` session that unblocks both installer fixes; then the hub naming; then the +2026-08-09 batch (R-279 … R-292) which is still untriaged against everything since. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index fdc5bfd..2a40115 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -508,7 +508,7 @@ applied.** The one that matters: Scenario A **fails against today's tree** with | **R-277** | **Three hub surfaces jointly present a HEALTHY off-site tier as an absent one — and it produced a wrong operator statement during this run.** For demo-hp on 2026-08-09 the box was pushing off-site daily without a gap (18 restic snapshots, `last_status: ok`), yet: (a) the customer page's Backup panel read `Snapshots 0 / Repo Size 0 MB / Integrity Unknown` — it renders the **local disk tier**, while the healthy `offsite` object sits **in the same report** unrendered on that panel; (b) the Offsite page read `0.0 GB` — true, but a 162 KB repo rounds to nothing; (c) a stale `offsite_delivery_stuck` event from **2026-08-07 10:19** (not recurring) reads as current state. **Three independent surfaces agreeing on a wrong picture is how a working backup gets "fixed".** It did exactly that here: the rehearsal reported a fleet-wide off-site outage to the operator and had to retract it. **Note the true half:** demo-felhom IS genuinely stuck (`offsite.state=needs_credential`, no run has ever succeeded) → **R-278** | **READY (S) — NEW 2026-08-09** | — | Render the offsite object on the offsite row; show bytes not rounded GB; distinguish a live alarm from event history | CC | | **R-278** | **demo-felhom's off-site tier has never completed a run and has been stuck for six days.** `offsite.state=needs_credential` since the 2026-08-03 guest rebuild; the hub's own alarm reads *"enabled + escrowed but no run has EVER succeeded"*; the controller's `offsite-credential-retry` job runs every 5 minutes and completes in 0 s, doing nothing. R-193's fix (the recovery SCREEN, controller 0.200.0) is present on the box, so the remedy exists — it just needs the customer-present ceremony that nobody has run, which is R-243's shape (*"a machine waiting for its recovery code can stop backing up off-site without alarming us"*) landing on a real box. **Contrast that makes it a defect and not a chore:** demo-hp, same rebuild, same day, recovered and has 18 snapshots | **CLOSED 2026-08-10 — protection RESTORED, and the recovery it waited for could never have worked** | — | Either the self-heal reconciler owns this shape end-to-end, or the box must say plainly on the dashboard that it is unprotected pending the recovery code **THE REMEDY THIS ROW ASSUMED WAS IMPOSSIBLE, and that is the finding.** The operator authorised the recovery ceremony on 2026-08-10; it was **not run**, because three measurements taken first showed it could not help. The box’s local key hashes to `c60c8bc737a6b7c6…`; the hub’s sealed escrow key hashes to **the same value**; and that key answers `Fatal: wrong password or no key found` against the repository. **Recovery would have returned a key the box already held and which was already proven not to open the store.** The repository was written under `48741892f0ef4d59…` (`host_escrow_superseded` id=4, superseded 2026-08-04 07:20:08) whose **`identity_blob` is NULL** — and the restic password lives ONLY in the identity bundle (`escrow/identity.go:39`, read by `escrow/recover.go:91` through `UnwrapIdentityBundle`), so it is unrecoverable by construction. Corroborated: the surviving K-escrow payload is **64 bytes**, a wrapped key, far too small to carry a bundle with a password. **The same shape the register already records for demo-hp** (*"its key sits in superseded row id 3 with identity_blob NULL … four hours before v0.93.0 fixed the retention"*). **What was done instead, on the operator’s stated fallback:** the orphan reset, through the customer’s own card — the old store **moved aside, never deleted**, to `/home/felhom-repo.orphaned-20260810` (1.2 GB); a fresh repository initialised under the current key; `offbox_repo_reset` audited hub-side at 08:06:31. **PROVEN rather than assumed:** `last_status: ok`, `last_success: 2026-08-10T08:07:33Z`, 10 s — and the snapshot’s CONTENTS listed, not just its count: `opengist/compose/{.felhom.yml,app.yaml,docker-compose.yml}`, `manifest.json`, and `volume-dumps/opengist_opengist_data.tar`. A week without off-site protection ends here | CC | | **R-279** | **There is no operator-triggerable off-site backup.** The only route to `POST /backup/offbox/run` is the customer's own dashboard session; `signed_jobs` carries opaque operator-SIGNED blobs and the hub holds no signing key. This cost the rehearsal a stop: preparing the run needed one off-site push and there was no operator path to it. Sibling of **R-177** (no operator-triggerable fill check) | **READY (XS) — NEW 2026-08-09** | — | Same shape as R-177; solve both together | CC | -| **R-280** | **RANK 1 — after a reinstall the data drive cannot be re-attached through ANY dashboard route, and the restore page promises it is "two clicks".** Measured on the rebuilt demo-hp, 2026-08-09. The restore page diagnoses the situation perfectly and then sends the customer to an empty page: *„Előbb csatold vissza az adatmeghajtót. A mentéseid megvannak, és a meghajtók is megvannak — újratelepítés után viszont a gép még nem ismeri őket, ezért most nincs hová visszaállítani. **Ez két kattintás:** Tárhely → Meghajtók, »Meglévő meghajtó csatolása«."* **It is not two clicks; it is zero possible clicks.** `GET /api/disks/candidates` → `{"initialize":[],"attach":[]}`, so both wizards render an empty selector, and `Tárhely → Meghajtók` reads „Nincs regisztrált adattároló" with an empty unregistered list. **The agent is not at fault** — `GET /api/disks` returns the NVMe in full (1.0 TB, SMART PASSED, `mount_path:/mnt/nvme-1tb`, `guest_attached:false`), so the channel and enumeration work. **ROOT CAUSE:** `handleDiskCandidates` builds both lists from `ListCandidateDisks`, the UNCLAIMED-disk scan; demo-hp's NVMe is deliberately BOTH the user-data drive and the `felhom-backup` target (`operations/nodes.md`), so it is claimed and never offered. That filter is **correct for `initialize`** (never offer to format a disk in use — `/storage/init` even says so: *„Rendszer- és biztonsági-mentés meghajtók itt nem jelennek meg — azok védettek"*) and **over-broad for `attach`**, which is non-destructive by definition and whose own page says *„A meghajtón lévő adatok nem törlődnek — a csatolás csak elérhetővé teszi azokat."* **It cascades:** no store → Calibre-Web's install page degrades to *„Nincs regisztrált adattároló — adja meg kézzel az útvonalat"* and demands a hand-typed `E-könyvtár útvonal`; no app → the restore rows read „Nincs telepítve". **THE ESCAPE HATCH WORKS AND NO CUSTOMER COULD FIND IT:** `POST /settings/storage/add` with `storage_path=/mnt/sys_drive` succeeded first try (*„Adattároló sikeresen hozzáadva"*) — and `/mnt/sys_drive` is an internal path, the very one registered before the wipe. Once registered, everything unblocked and the deploy form became a proper picker (*„Tárhely (sys_drive) — 64.2 GB szabad"*). **This is R-220's successor:** R-220 was closed as "drives unenrollable after a rebuild — fixed"; enumeration is fixed, OFFERING is not | **READY (M) — NEW 2026-08-09** | — | Populate `attach` from mounted-but-unregistered filesystems rather than from the unclaimed-DISK scan; and never print "two clicks" without asserting the destination is non-empty | CC | +| **R-280** | **RANK 1 — after a reinstall the data drive cannot be re-attached through ANY dashboard route, and the restore page promises it is "two clicks".** Measured on the rebuilt demo-hp, 2026-08-09. The restore page diagnoses the situation perfectly and then sends the customer to an empty page: *„Előbb csatold vissza az adatmeghajtót. A mentéseid megvannak, és a meghajtók is megvannak — újratelepítés után viszont a gép még nem ismeri őket, ezért most nincs hová visszaállítani. **Ez két kattintás:** Tárhely → Meghajtók, »Meglévő meghajtó csatolása«."* **It is not two clicks; it is zero possible clicks.** `GET /api/disks/candidates` → `{"initialize":[],"attach":[]}`, so both wizards render an empty selector, and `Tárhely → Meghajtók` reads „Nincs regisztrált adattároló" with an empty unregistered list. **The agent is not at fault** — `GET /api/disks` returns the NVMe in full (1.0 TB, SMART PASSED, `mount_path:/mnt/nvme-1tb`, `guest_attached:false`), so the channel and enumeration work. **ROOT CAUSE:** `handleDiskCandidates` builds both lists from `ListCandidateDisks`, the UNCLAIMED-disk scan; demo-hp's NVMe is deliberately BOTH the user-data drive and the `felhom-backup` target (`operations/nodes.md`), so it is claimed and never offered. That filter is **correct for `initialize`** (never offer to format a disk in use — `/storage/init` even says so: *„Rendszer- és biztonsági-mentés meghajtók itt nem jelennek meg — azok védettek"*) and **over-broad for `attach`**, which is non-destructive by definition and whose own page says *„A meghajtón lévő adatok nem törlődnek — a csatolás csak elérhetővé teszi azokat."* **It cascades:** no store → Calibre-Web's install page degrades to *„Nincs regisztrált adattároló — adja meg kézzel az útvonalat"* and demands a hand-typed `E-könyvtár útvonal`; no app → the restore rows read „Nincs telepítve". **THE ESCAPE HATCH WORKS AND NO CUSTOMER COULD FIND IT:** `POST /settings/storage/add` with `storage_path=/mnt/sys_drive` succeeded first try (*„Adattároló sikeresen hozzáadva"*) — and `/mnt/sys_drive` is an internal path, the very one registered before the wipe. Once registered, everything unblocked and the deploy form became a proper picker (*„Tárhely (sys_drive) — 64.2 GB szabad"*). **This is R-220's successor:** R-220 was closed as "drives unenrollable after a rebuild — fixed"; enumeration is fixed, OFFERING is not | **CLOSED — controller v0.211.0, delivered via golden 0.211.0 (vouched 2026-08-10)** | — | Populate `attach` from mounted-but-unregistered filesystems rather than from the unclaimed-DISK scan; and never print "two clicks" without asserting the destination is non-empty | CC | | **R-281** | ~~**The hub said NOTHING through an entire reinstall — and the tripwire for a sealed-backup unseal did not fire on a real unseal.**~~ **WITHDRAWN 2026-08-09 — THE FINDING WAS AN ARTEFACT OF MY OWN MEASUREMENT, AND IT WAS WRONG IN BOTH DIRECTIONS.** The operator's mailbox settled it: the hub fired **twenty events** on 2026-08-09, and `escrow_blob_served` **DID** fire — 10:19:41 UTC / **12:19 CEST**, eight minutes before the verified restore. **Cause of the false reading, ESTABLISHED (not guessed):** the P7 query copied `/data/hub.db` **without `hub.db-wal`**. The hub runs SQLite in WAL mode (R-172), so every write since the last checkpoint was invisible. **The signature is an exact match:** P7 reported *"2 events all day, newest `db_dump_completed` 00:30:07"*, and the number of rows on 08-09 at or before 00:30:07 is **exactly 2**. **The two obvious alternatives were TESTED AND REFUTED**, not waved away: a **timezone offset** — all nine mailbox stamps equal the hub's UTC + 2 h exactly (`escrow_blob_served` 10:19→12:19, `host_down` 09:28→11:28, and seven more), so the window was right; and a **wrong customer key or wrong store** — the same table and key return the correct rows now. A live re-run cannot reproduce the fault because the WAL has since been checkpointed; the case rests on the command text plus the 2-of-2 count signature, and that is stated rather than dressed up as a reproduction. **This is a trap this project has already documented** — `operations/nodes.md` says copying `hub.db` alone is *"valid but stale … the worst failure shape"* — and I had avoided it correctly earlier in the same session before hitting it. **Split out: → R-285** (the real, opposite defect) and **→ R-286** (the measurement lesson) | **WITHDRAWN 2026-08-09** | — | Superseded by R-285/R-286 | CC | | **R-282** | **One secret, three different Hungarian names, and the email sends the customer to a page their box is not showing.** Sending it from the hub is „**Visszaállító** kód küldése"; the email that arrives is subject „Jelszó-**visszaállítási** kód", body „**Visszaállító** kód: …", and it instructs *„Add meg a vezérlőpult »**Elfelejtett jelszó**« oldalán"*; the page the box actually serves is „A szerver **beállítása**" asking for a „**Beállító** kód". **A rebuilt box shows a SETUP page and the hub can only send a RESET mail** (because hub-side the customer is still `claimed_at 2026-07-21`), so the instruction names a route that does not exist on screen. **It does work if you ignore the instructions** — the reset code was accepted on the setup page (302 + session), so this is naming, not function. **It cost this session real time and one wasted code:** the operator supplied a 3-word Hungarian code believing it was the recovery code, because the hub calls the claim code „Visszaállító kód" and the ESCROW code is also „Visszaállító kód" — the only reliable discriminator is length (claim = 3 Hungarian words; recovery = **10** EFF-list words, and the recovery screen does say „(tíz szó)") | **READY (S) — NEW 2026-08-09** | — | Pick one name per secret and use it on all three surfaces; make the mail's page reference match what a rebuilt box actually shows | CC | | **R-283** | **After a rebuild the hub says "Claimed 18d ago" while the box serves its first-run setup page.** `customer_claims` for demo-hp still read `claimed_at 2026-07-21 16:29:25`, `generation 2`, `issued_at 2026-08-03` while the freshly provisioned guest — whose `settings.json` is new — correctly showed „A szerver beállítása". The two sides never reconcile: the hub's claim state survives a guest rebuild and the box's does not. Consequences: the operator's screen says the box is claimed when it is not, a resend produces a RESET code instead of a SETUP code (→ **R-282**), and any previously issued code fails with *„Hibás vagy lejárt kód"* — a message that is technically true and tells the customer nothing about the real cause, namely their own reinstall. Mirror image of **R-214/R-235** (an already-paired box still told to pair itself) | **READY (S) — NEW 2026-08-09** | — | Let a report from a box carrying no claim state clear the hub's, or show both sides on the operator page | CC | @@ -522,7 +522,7 @@ applied.** The one that matters: Scenario A **fails against today's tree** with | **R-291** | **CI's installability assertion is now BOUNDED by a retention number, and the narrowing is recorded here so it can be widened deliberately rather than discovered.** `check-published-versions.py` demanded that **every** `v` tag still be downloadable while the registry demonstrably does not retain every version — two sensible rules that cannot both hold, which is why CI went red at a commit whose own run had been green the day before, and would have gone red again at the next publish. **The fix couples them:** `felhom-agent/scripts/retention-policy.json` is THE number (`generic_versions_kept: 10`) and the check reads it. **WHAT CI NO LONGER COVERS, stated plainly: a released version older than the retention window is no longer asserted downloadable.** Its git TAG and its config tree are still asserted — only the binary's presence is dropped — and the check **prints the dropped versions on every run**, so the narrowing cannot go quiet. Controls run: widened to 11 the evicted version re-enters and convicts (exit 1); the policy file removed gives INCONCLUSIVE (exit 2), never silently unbounded. **The number is an OBSERVED state, not a located ruling** (R-287) and the file says so. **The better bound, recorded rather than built:** the hub's vouched `min_agent` floor — nothing can install an agent below it, so a sub-floor version being un-downloadable costs nothing real; it needs the gate to read the hub, which is network it does not have today | **READY (S) — NEW 2026-08-09** | R-287 | ~~Widen or replace the number when the deleter is established~~ — **CONDITION RELEASED 2026-08-10: the deleter IS established (R-287), so the operator is no longer blocked on establishing what was already written down.** The number can now be confirmed or replaced on its merits. The better bound remains the vouched `min_agent` floor | CC | | **R-292** | **The artifact-save flash conflates three different facts, and a failing test found it rather than a reading.** `artifact_sha_invalid` reads *"the Gitea sha lookup failed (version missing / Gitea unreachable) or the manually-entered sha is invalid"* — three causes, one message, and the operator acts differently on each. It surfaced because scenario E of the new installability gate kept reporting `artifact_sha_invalid` where it expected `artifact_unverifiable`: `resolveArtifactSHA` ran first and swallowed the distinction. **Worked around in v0.102.0 by ORDERING** — the installability probes now run before the sha resolution, so an unreachable registry is reported as unreachable — **but the underlying message is untouched and still conflates on its own paths** | **READY (XS) — NEW 2026-08-09** | — | Split it into "version not found", "registry unreachable" and "invalid sha" | CC | | **R-293** | **CENSUS, 2026-08-10 — no machine that is not ours can be in the state that cost demo-felhom its history, and here is the whole population.** Read-only against the hub store, with a control run first (the query returned *present (572 bytes)* for a host known to have material and *absent (NULL)* for one known not to — both cases from the two demo boxes, so the instrument was shown to distinguish the states before it was trusted). **The hub knows three hosts.** `demo-felhom-8363b5` and `demo-hp-bb76ea` each hold one superseded escrow with **`identity_blob` ABSENT**, superseded **07:20:08** and **07:15:36** on 2026-08-04 — both **before** the retention fix was in force, pinned at **11:11:37Z** from the hub's own first post-fix escrow row (a date-only comparison mislabels these as "after" and was corrected). `drill-r50-0a4f9a` has no supersession. **`peti-felhom` — the tester's machine — has NO host row and NO escrow at all**, and neither does `david`; the orphan check found no escrow row pointing at an unknown host. **So the answer is: no, not today, and not tomorrow either** — any future enrolment escrows under the fixed code. **What this does NOT claim:** that a retained blob has ever been *unwrapped* on a superseded row. Retention is proven; the recovery FROM a superseded row is still unexercised | **CLOSED-INFORMATIONAL 2026-08-10** | — | The machine was not contacted; only the hub's records were read | CC | -| **R-294** | **The orphan card promises restorability that the box rendering it cannot evaluate — specified, not implemented.** The card says the set-aside copies *"a hozzá tartozó helyreállítási kóddal később visszaállítható lehet"* (`controller/internal/web/templates/backups_remote.html:101`). **The discriminator lives in the hub** (`host_escrow_superseded.identity_blob`, `hub/internal/store/store.go:393`); **the box caches only `HubEscrowIdentityPresent`** (`controller/internal/settings/settings.go:71`), which describes the CURRENT escrow, not a superseded one; and **no field on the report or ACK wire carries superseded-blob retention**. So the renderer cannot tell which case the customer is in — **a conditional promise the system cannot evaluate is the same defect as an unconditional false one.** Replacement Hungarian copy, the surfaces at `file:line`, and the render tests that should pin it: **`documentation/design/SPEC-orphan-card-copy-2026-08-10.md`**. **Deliberately NOT implemented:** it lands in the controller, and a controller release is undelivered until a golden carries it (R-242) — one bake, one approval. **It should ship with the next controller change so one bake covers both**, and the spec says so. Parent: **R-202** | **READY (S) — NEW 2026-08-10** | R-202 | Implement with the next controller release, not on its own | CC | +| **R-294** | **The orphan card promises restorability that the box rendering it cannot evaluate — specified, not implemented.** The card says the set-aside copies *"a hozzá tartozó helyreállítási kóddal később visszaállítható lehet"* (`controller/internal/web/templates/backups_remote.html:101`). **The discriminator lives in the hub** (`host_escrow_superseded.identity_blob`, `hub/internal/store/store.go:393`); **the box caches only `HubEscrowIdentityPresent`** (`controller/internal/settings/settings.go:71`), which describes the CURRENT escrow, not a superseded one; and **no field on the report or ACK wire carries superseded-blob retention**. So the renderer cannot tell which case the customer is in — **a conditional promise the system cannot evaluate is the same defect as an unconditional false one.** Replacement Hungarian copy, the surfaces at `file:line`, and the render tests that should pin it: **`documentation/design/SPEC-orphan-card-copy-2026-08-10.md`**. **Deliberately NOT implemented:** it lands in the controller, and a controller release is undelivered until a golden carries it (R-242) — one bake, one approval. **It should ship with the next controller change so one bake covers both**, and the spec says so. Parent: **R-202** | **CLOSED — controller v0.211.0; see R-299 for the sentence it missed** | R-202 | Implement with the next controller release, not on its own | CC | **Explicitly still open, untouched by this session:** R-246, R-255, R-256, R-257, R-261, R-262, R-263, **R-264** (the twenty-one undecided facts — a design session of its own), R-240, R-243,