From ceac5e0deb0c1ad51d7705b6800086fc9761fac7 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Fri, 7 Aug 2026 12:59:52 +0200 Subject: [PATCH] STATUS.md rebuilt (258 -> 85); R-245 re-filed as decided; R-246/247/248 filed STATUS.md REBUILT FROM THE REGISTER, not trimmed. Its own header says one screen; it had reached 258 lines, having been 83 four days ago. The three named defects, all fixed: 1. the "waiting on you" list asked the operator to decide the RECOVERY SCREEN, built and shipped 2026-08-05, and to approve an orphaned-backup deletion the register records as DONE the same day; 2. a stray line reading only "- **Nothing.**" sat mid-list; 3. the DooPlex infrastructure work was mixed in with the product's. Infrastructure is now under ITS OWN HEADING rather than dropped, and the reason is stated on the page: these are real asks that need the operator, but they concern the machine this is built on, not what a customer receives. Dropping them would lose real work; mixing them is why the page stopped being readable. The 100-line "what shipped recently" log is gone. That is what the per-repo CHANGELOGs and the register are for, and restating it here is what made the page grow back. R-245 RE-FILED as a decision taken, not a question pending. It sat as WAITING-ON-OPERATOR for a day with nothing actually pending - it was settled on 2026-08-07. It keeps the whole reasoning and now carries the condition that would REOPEN it, which the reasoning already named: QUOTA, old set-aside history blocking new backups. A condition, not a calendar. AUDIT OF EVERY WAITING-ON-OPERATOR ROW, parsing the state column exactly rather than grepping for the phrase (which over-matches rows that merely mention it): exactly ONE row carried it - R-245 - and it was a settled decision. So zero rows were genuinely waiting, and the drift was caught while it was still a single row. R-246/R-247/R-248 file the read-only stale-blob spike's findings. R-242 updated: it recurred within a day, and shape (b) is now built - with the vouch half explicitly still open on that row rather than being papered over. --- STATUS.md | 307 ++++++---------------------- documentation/backlog/OPEN-ITEMS.md | 8 +- 2 files changed, 73 insertions(+), 242 deletions(-) diff --git a/STATUS.md b/STATUS.md index c55f43a..e6497d4 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,258 +1,85 @@ # STATUS — what works, what's broken, what's next -**Updated 2026-08-07.** +**Updated 2026-08-08.** -> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority on open work; this -> page restates part of it in plain words, and **nothing may exist only here**. **Not `CONTEXT.md`**, -> which is technical state written for Claude Code — keep the two separate. **Maintenance:** update -> at the end of every session in which something shipped, broke, or was decided. One screen; cut -> items rather than extend it. +> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates +> part of it in plain words, and **nothing may exist only here**. Not `CONTEXT.md`, which is technical +> state written for Claude Code. **Items, not paragraphs. One screen.** If it does not fit, something +> belongs in the register instead. +> +> *Rebuilt from the register on 2026-08-08. It had reached 258 lines; its "waiting on you" list asked +> for two things already shipped and carried a stray line reading only "Nothing."; and it mixed the +> DooPlex infrastructure work in with the product. The old "what shipped recently" log — 100 lines of +> it — is what the per-repo `CHANGELOG.md` files and the register are for, and is not restated here.* -## What works right now +## What works -A blank machine boots the Felhom disc, installs itself unattended, and is claimed by the customer, -who sets their own password. They install apps from a catalogue of fifty-three, share files over the -home network, and open apps from a launcher or a shared link. Backups run on their own to three -places — the machine's drive, a second drive, and an encrypted off-site copy. +A blank machine boots the Felhom disc, installs itself unattended, and is claimed by the customer, who +sets their own password. They install apps from a catalogue of fifty-three, share files over the home +network, and open apps from a launcher or a shared link. Backups run on their own to three places — the +machine's drive, a second drive, and an encrypted off-site copy. -**The backup promise is proved again. The recovery JOURNEY still is not — but the fault that broke it -last night is now fixed.** On 6–7 August we built a brand-new machine from the published disc, gave it -three marked files, soaked it through a full night, destroyed it and rebuilt it. **The files came back -byte for byte identical**, all three, including one with Hungarian accents in its name. **But the -household had no way to ask for them:** the screen that takes their recovery code had switched itself -off, and the backup page offered to make them a *new* code instead. - -**7 August — we asked which of two things was wrong, and the answer reversed the repair.** The obvious -reading was that the screen's rule was too narrow. It was not. **The screen was telling the truth** — -there really was nothing openable with the key the machine held, **because the machine had made that -key itself**, on top of the sealed package we were already keeping for it. And it knew: it wrote *"the -sealed package does not cover the current key"* into its own log **thirty-five minutes before the -household looked**, then threw the answer away. Mending the screen would have hidden a machine quietly -making its own backups unopenable. *(R-241)* - -**Fixed the same day.** The machine **stops making its own key** while we hold a package for it — the -repair that prevents the situation rather than tidying up after it. **The comparison it was already -making now decides whether to offer help**, instead of two indirect guesses that have each been wrong -in opposite directions. And **giving up the old backups became a finishable thing**: a **14-day -countdown** you can see and change your mind about, at the end of which the old backups *and* their -sealed package go together — so the question stops coming back because there is nothing left to ask -about, not because something is suppressing it. On the way past, the *"create a new recovery code"* -button is now **unavailable** while a recovery is outstanding, and the confirmation says plainly that -the old backups are deleted **on a date** rather than merely set aside. - -**Two of our own mistakes were caught by tests rather than by reading the code**, which is the point of -having them: one would have warned a household that had switched off-site backups off, and one would -have brought back the very fault we were fixing. - -**Still open, on purpose:** a machine held waiting for its code raises no alarm to *us* *(R-243)*; -nothing enforces that a release reaches a new machine *(R-242)*; and we deliberately did **not** build -automatic abandonment after 30 days *(R-245 — the reasoning is written down)*. - -**Today's machines now get last week's fixes — approved 7 August.** *(R-239 — closed.)* +**The backup promise is proved.** A machine has been destroyed on purpose and its files came back byte +for byte identical — three separate times, including a filename with Hungarian accents. ## What's broken -- **A backup could report „✓ Rendben" while quietly leaving out an app the customer had just - chosen — FIXED 6 August.** Two things were wrong and only one had been guessed at. The machine's - own rule said *"a warning beside a success is read as a success"*, and it applied it to an app - missing a *folder* but not to an app left out *entirely*; that now counts too. And the real cause of - the case we saw: pressing „Távoli mentés most" while a backup was already running answered - „elindult" and then showed the **previous** run's green tick — so the customer read it as covering - their new app. It did not, and the restore refused minutes later. The button now says plainly that - it did not start anything. +- **A household still cannot get their own data back unaided.** Every individual link now works; no + single walk has completed end to end without someone stepping in. *(R-201)* +- **The machine's own screen keeps telling an already-paired box to pair itself** — 25 minutes after it + was paired, on a screen that promises it refreshes itself. *(R-214, R-235)* +- **A rebuilt machine cannot create a new recovery code at all.** *(R-221)* +- **A backup that covered nothing still calls itself „Sikeres".** The state is honest; the word is not. + *(R-240)* +- **A machine waiting for its recovery code raises no alarm to us.** It quietly stops making off-site + backups, and three separate safety nets each correctly decide it is not their business. The + household can see it; we cannot. *(R-243)* +- **The card offering to reopen set-aside backups promises more than we can deliver** — we keep the old + sealed package, but nothing can open it. *(R-202)* +- **Deleting a customer leaves rows behind** on every test machine ever torn down, while reporting a + clean teardown. No secrets involved, but it accumulates with each walk. *(R-244)* +- **Putting restored files back where they belong is still a manual step.** *(R-213)* -- **~~A machine installed today still gets the older in-house service~~ — RESOLVED 7 August.** The - newer service is packaged and approved, so a newly installed machine now gets it without anyone - touching the machine. *(R-216, R-223, R-224, R-239)* -- **Three things a rebuilt machine still cannot do by itself.** Its owner cannot re-attach their own - drives, so no app can be put back on its data; it cannot create a new recovery code at all; and the - screen at the machine itself never stops showing a stale pairing code. Each is understood, measured - and written down — none is fixed yet. *(R-220, R-221, R-214)* -- **~~Rebuilding a machine throws away its off-site backup history~~ — the CAUSE is fixed (7 August).** - A rebuilt machine used to invent a new encryption key over the top of the sealed package we hold for - it. It no longer does: while we hold a package, it waits for the household's recovery code instead. - The old key is kept and a changed key still raises an alarm the same day. *(R-193, R-198, R-241)* -- **The kept older backups cannot be opened — by anyone.** We keep the previous sealed package and - there is no way to open it. The screens now say exactly that and stop. **One place still promises - otherwise**: the older-backups card says they "may be restorable later with the matching code", - which is not true today. *(R-222, R-202)* -- **The off-site copy can be erased by the machine that made it.** A daily snapshot is armed as a - stopgap. *(R-95, R-87)* +## Found today -## Can a household get their data back on their own? Asked again today — still no, but nearer - -We built a **brand-new machine** from the published disc, gave it three marked files, destroyed it — -guest and both drives, as a hardware loss would — and tried to get them back the way a household -would. *(R-201, the re-walk)* - -**The files came back perfectly.** All three, byte for byte, including a 12 MB file and one whose -Hungarian accented filename came back **letter-for-letter identical**. Out of the pre-destruction -backup, in **23 seconds**, through the customer's own restore screen. - -**And much of the journey now works.** The machine showed the recovery screen **without being asked**, -told the customer what was waiting and when it was sealed, said plainly that nobody can replace a lost -code, and accepted the real code first time. The emailed claim code worked first try. - -**But it still needed us twice**, and a household has neither hand: - -- **The machine never picks up its own storage connection.** Our hub hands it over and says "the box - will collect this on its next cycle" — the cycle came and went and it did not. Nothing the customer - can click fixes it; it took a command inside the machine. *(R-218 — we had recorded this as fixed; - only half of it was)* -- **A rebuilt machine still cannot re-attach its own drives** — and without them no app can be put - back, so the restore screen stays empty. *(R-220)* - -**Two dead ends, down from four.** The verdict stays **FAILED** until a walk needs us zero times. - -**One thing to decide.** A machine installed today still gets the older software — **the fixes are -built and published but not approved for new machines**. We installed them by hand for this test. So -this proves the journey works on the fixed build; it does **not** prove a customer would receive it. - -## What we fixed this morning, and what it did not fix - -Overnight we tried to break the recovery journey with eleven faults and then left the machine alone -for a full cycle. Nothing lost a byte; the set-aside backups really were kept; the alarm fired and -cleared itself. What it found was that **the machine blamed the customer for failures that were not -theirs** — our hub being unreachable came back as "your recovery code is wrong", in three hundredths -of a second, without the machine even trying. **That is fixed and deployed** *(R-224, R-226)*, along -with three smaller truths: "0 snapshots" where the answer is "we have not looked yet" *(R-225)*, a raw -English error mid-recovery *(R-227)*, and set-aside backups that had become invisible *(R-228)*. - -**None of that shortened the journey**, which is why today's re-walk above still says FAILED — the two -remaining dead ends are different ones. - -## What shipped recently - -- **2026-08-06 (later)** — **The page that says which machine may be wrecked had one wrong sentence, - and it has been checked against the machine rather than corrected from memory.** It claimed the - spare HP box keeps no off-site copy, and used that as the reason to aim risky backup tests at the - other one. The HP box **does** keep an off-site copy — two of them, the newest from 4 August, in - its own space on the off-site server. The sentence was true when it was written and went out of - date on 23 July. It matters because a belief that a machine holds nothing is how a destructive test - lands on one that does. Also today: the notes the assistant keeps had three statements that were - simply untrue (including that the demo boxes were still away) — those three are fixed, the rest are - now flagged automatically rather than rewritten by hand. *(R-229, R-230)* -- **2026-08-06** — **The assistant's own working notes were on one machine with no copy anywhere.** - Everything the assistant has learned about this system over months — where things live, which traps - cost us an incident — sat in a folder on the build machine that no backup touched and no repository - held. It is now included in that machine's nightly backup. **Two things you should know before - treating that as solved:** the copy lands on the **same physical disk** as the original, so it - survives a mistake but not a dead drive, and the build machine's backups have **no off-site copy at - all**. **A full survey of that machine's backup, done the same day, confirmed both and found two - more things worth knowing.** The good news first: it has run every night without missing a set, and - we pulled a file back out of it and checked it matched the original exactly — the first time that - has ever been demonstrated. The rest: **if a backup ever fails, nobody is told** — the alert was - configured but never given anywhere to send to — and **every copy it makes stays inside that one - box**, so it survives any single disk dying but not the room. Nothing was changed; the survey was - read-only and the decisions are yours. Also tidied the same day: a quarter of those notes had become unreachable — filed but listed - nowhere, so nothing would ever read them — and the instructions the assistant reads at the start of - every session were cut roughly in half, with a check added so they cannot quietly grow back. Nothing - was deleted. *(R-229, R-230, R-231)* -- **2026-08-05 (late)** — **New machines now get current software again — and the disc image had been - quietly un-updatable for days.** Approving the newer in-house service turned out to be impossible on - its own: the system correctly refuses to publish a set where the pre-built machine image is older - than the software the fleet already runs, and that image had been behind since late July. **So the - image was rebuilt and both were published together.** A machine installed from now on lands on - current software and can open a recovery package on day one. *(R-223)* - -- **2026-08-05 (evening)** — **Four sentences where there was one, and none of them blames you.** - "We did not accept your recovery code" used to appear when the code was wrong, when the machine - could not ask, when the store could not be read, and when the customer held the code for an older - backup we still keep. Only the first is the customer's doing. Also today: succeeding at recovery no - longer switches off the machine's own request for the thing it still needs; the screen now finishes - the job and shows what is in the backups instead of promising a list it could never produce; and the - recovery page can no longer be reached on a machine that never had backups. - *(R-216, R-217, R-218, R-219, R-222, R-215)* - -- **2026-08-05** — **The recovery screen: a customer whose machine was rebuilt is now told, and shown - how.** Until today they had everything needed to get their data back and no way to find out — the - only route was a command line. The screen unlocks the backups and lists what is in them; it does - **not** restore anything, because unlocking and restoring are two different decisions and mixing - them would turn one clear moment into a wizard. Three ways out, none of them a dismiss button — and - the „most nem" option keeps the route to the data permanently visible in the backups area, because - a notice someone clicks past once is a notice that never happened. *(R-193)* - -- **2026-08-05** — **The orphaned backups are deleted — and the list you were given was wrong, which - is why you were asked again.** You had approved "about 1.2 GB in two set-aside stores". Measured - before touching anything: there were **three** set-aside stores totalling **~1.45 GB** — and the - thing that was exactly 1.2 GB was demo-felhom's **live** store. Matching on the size would have - deleted a working backup. With the corrected list confirmed, all three were removed and both live - stores left alone; a real off-site backup ran successfully straight afterwards to prove nothing - working had been caught. *(R-212)* - -- **2026-08-05** — **A rebuilt machine now asks for its storage credential, and the hub gives it back.** - The machine says plainly what it needs — it can tell it has been rebuilt, because its data area is - empty *and* the hub is holding a sealed recovery package for it — instead of leaving the hub to guess - from a silence that has four possible meanings. The hub waits long enough to be sure it is not a - restart, then **re-uses the credential it already holds** before creating a new one at the storage - provider. **The recovery ceremony stays manual, deliberately:** a credential can be replaced, your - recovery code cannot. *(R-204 item 4, R-193)* - -- **2026-08-05** — **The DooPlex server's disk is out of danger: 86% full → 54%, and the storage layer - is unstuck.** The cause was leftover working data from building our own software — 157 GB of it, - growing about 5 GB a day, which nothing was allowed to delete. **148 GB came back in 86 seconds.** - The storage layer had already stopped accepting new copies of any volume onto that disk; that is - fixed the same day. A **30 GB ceiling** is now in place and was **proved to work by deliberately - overfilling it and watching it evict** — not by assuming the setting took. Two things that failed - quietly around it: the warning meant to catch exactly this **could never fire** (fixed and proved), - and the weekly cleanup is still forbidden from touching the thing that grows (next session). - *(R-205 … R-211)* -- **2026-08-05** — **The build data now lives on the second SSD, moved with nothing lost.** All 345 - images, every saved volume and both development databases came through identical — checked before - the original was touched and again afterwards, and confirmed by running a real build on the moved - copy. **The server's own services never went down:** Gitea, the registry, the hub and the backup - system run on a separate system and stayed up throughout. The second SSD now also **reserves 80 GB** - for this, so the storage layer can no longer quietly claim the space and repeat what happened to the - first disk. The old copy is kept as the way back until the machine next restarts. *(R-209)* -- **2026-08-05** — **Three of the four recovery crutches removed.** The reset code works first time; - a credential re-issue no longer blocks off-site backups on a healthy machine; and the default - restore no longer quietly returns the wrong thing. *(R-204 items 1–3, R-196)* -- **2026-08-04 (night)** — **The drill PASSED**, end to end, on real hardware. *(R-201)* -- **2026-08-04** — The folder-left-out-of-the-backup problem fixed both halves: the app and its backup - look in the same directory, and a backup that misses a folder marked essential reports *incomplete* - instead of success. *(R-203)* -- **2026-08-04** — The hub now keeps the off-site backup key when a machine re-seals, instead of only - the whole-machine one, and a machine can fetch its own sealed package back. *(R-198, R-199)* -- **2026-08-04** — The daily false alarm about David is gone; a vanished permission now repairs itself - **and says it had to**; the weekly off-site backup stopped reporting failure after a successful - upload. *(R-195, R-190, R-191)* +- **A leftover flag has been silently switching off the new recovery detection on one demo machine + since 4 August.** A Re-issue during the recovery drill set it, using code we removed the next day. + **The flag is wrong** — the sealed package does cover the key that machine is using, and the two + fingerprints match exactly. Because of it the hub withholds a figure the machine needs, so the + machine tells its owner *"create a new recovery code"* — the one act that would put their old backups + beyond reach. **A freshly installed machine cannot reach this state**, because nothing has set that + flag since 5 August. *(R-246, R-247, R-248)* ## What we're working on -- **Next: the retention proof.** The hub keeping the old sealed key when a machine re-seals is the one - remaining link that has **never run outside a test**. Proving it needs a second deliberate wipe on - the spare demo machine, and it is its own procedure. *(R-198)* -- **Then:** the orphaned-backup deletion you asked for (below), and the off-site copy the machine can - still erase. *(R-193, R-95)* +- **The walk that either finishes the arc or says why not.** It needs the base image current (done + today) and a machine whose recovery detection is not silently switched off (answered today). *(R-201)* +- **Proving the hub really keeps the old sealed key** when a machine re-seals. Never run outside a + test; needs a second deliberate wipe and its own session. *(R-198)* ## Waiting on you +- **Approve the new base image**, so every future installation carries this week's fixes. *(R-239, + R-242 — presented separately in this session.)* -- **One thing to read after the machine next restarts — and nothing to do until then.** You told me - not to restart DooPlex, so I did not, and the move to the second SSD has therefore never been - through a restart. It works right now and nothing was lost, but a restart is the one test that - matters for this kind of change, and it has not happened. **I made it check itself:** whenever the - machine next starts, for any reason, it writes a plain PASS or FAIL line to - `/var/log/felhom-store-postboot-check.log`. **If it says PASS, the old copy can be deleted and 34 GB - comes back.** Until then I have deliberately kept that old copy, which is the only quick way back — - it is why the disk sits at 54% rather than lower. *(R-209a)* -- **One list to rule on: 193 old images that exist only on this machine.** 131 controller versions and - 62 hub versions are not in the registry, so they cannot be re-downloaded — all of them old - (controller up to 0.135.0, hub up to 0.57.0; everything newer is safely in the registry). **Nothing - was deleted.** Worth knowing before you spend time on it: they only account for about 27 GB against - 199 GB now free, so this is about clutter, not space. *(R-210)* -- **Nothing.** -- **The recovery screen you described has been priced, and it can be built.** A freshly installed - machine that finds a sealed package waiting should say so, offer a box for the recovery code, and - show what would come back before doing anything. One thing to weigh, deliberately not decided: that - screen is reachable by anyone with the household's dashboard password, and the preview reveals - backup dates and app names. *(R-193)* -- **The orphaned backups on the storage box — you said delete, and it is still owed.** About 1.2 GB - across the two demo machines, in set-aside stores nobody can open and nothing prunes. It wants its - own session rather than riding along with other work. *(R-193)* -- *(decided 4 Aug)* You chose **not** to keep a copy of the backup key on the Proxmox host, which - makes the customer's own recovery code the only route back from a rebuild. *(R-193)* -- **A job, not a decision: the hub password needs changing.** A diagnostic command printed it into a - session log; nothing suggests anyone else saw it. *(R-132)* -- **One small question, not urgent.** The automatic version check cannot see which version you have - told machines to install, only which ones exist. *(R-184)* +*Nothing else is pending. R-245 — whether an undecided household is auto-abandoned after 30 days — was +**settled on 7 August** (we do not build it, and the reasoning is recorded). It has been re-filed as a +decision taken rather than a question sitting in your queue.* + +## DooPlex infrastructure — separate from the product + +*Kept under its own heading rather than dropped: these are real asks that need you, but they concern +the machine all this is built on, not what a customer receives. Mixing them in is why the page stopped +being readable.* + +- **DooPlex's own backup makes every copy inside the same box, and nothing says when it fails.** + *(R-232)* +- **193 old images exist only on this machine** and cannot be re-downloaded. About 27 GB against 199 GB + free — clutter, not space. Nothing deleted. *(R-210)* +- **The hub password needs rotating** — a diagnostic printed it into a session log; nothing suggests + anyone else saw it. *(R-132)* +- **One thing to read after DooPlex next restarts.** The move to the second SSD has never been through + a reboot; it now writes PASS or FAIL to `/var/log/felhom-store-postboot-check.log` on every start. + On PASS, 34 GB comes back. *(R-209a)* +- **Backup scripts on DooPlex are unversioned host state.** *(R-231)* +- **Instruction-file follow-ups**, each needing a decision rather than an edit. *(R-229, R-230)* diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index ec190cc..02a00ff 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -169,12 +169,16 @@ the fault was real. Full observables: `tests/campaign11-evidence-2026-08-05/jour | **R-241** | **The credential self-heal, succeeding, locks the customer out of their own recovery.** Measured end to end on the final walk (2026-08-07, `tests/finalwalk-r201-2026-08-07/journal.md`). `OffsiteRecoveryOffer()` shows the recovery screen on exactly two conditions: **(a)** the box has **no** repository password — the pristine rebuilt shape — or **(b)** it has one but the inherited history will not open under it (`OffboxOrphaned()`). Overnight, unaided and exactly as designed, `offsiteheal` re-staged the one-time credential and the box's 5-minute retry **collected it and applied the tier**, writing a **fresh repository password** at 03:18Z. That makes **(a) false**. **(b)** is false too, because orphan detection only fires when a run actually tries the repository — and runs are blocked by `escrow_state: pending`. **The box therefore sits in the gap between the two conditions, and the gap is self-locking:** it cannot detect the orphan without running, cannot run without escrow, and cannot escrow without minting a NEW recovery code — which would orphan the history the customer's existing code protects. **What the customer sees:** `/` is „Indítópult" with no recovery pointer; `/recovery` **302s away**; `/backups/remote` offers „Helyreállítási kód **létrehozása**". **There is no field anywhere to enter the code they hold.** **And the operator's documented remedy also refuses** — `--recover-offsite-install` returns *„[REFUSED] a DIFFERENT repository password is already present… which history to keep is not a decision this command may take. Nothing written."*, which is correct and fail-closed and still a dead end. Recovery required moving the fresh key aside by hand and re-running the install: **three guest command lines**. **The two keys, measured:** on-disk `9b4a9a9d…` (self-heal) vs recovered-from-R `30ef574f…`. **THE DATA WAS NEVER AT RISK** — all three sentinels restored byte-identical once the right key was in place. **This is R-218's shape one level up:** that finding read *"succeeding at recovery stopped the box asking for what it still needed"*; here, succeeding at the credential self-heal stopped the box **offering** the recovery it still needed. The same walk proved the self-heal working unaided six hours earlier, and that success is what causes this. **Likely shape of the fix, not yet a decision:** the offer needs a third condition — a box holding a password it has never successfully used, while the hub holds a sealed package, is a recovery candidate — or the self-heal must not install a credential on a box whose escrow is still `pending` and whose hub blob is unconsumed. **Which of those is right is a design decision, deliberately not taken here.** **⚠ RULED 2026-08-07 by a read-only spike on the standing venue — `audits/SPIKE-r241-recovery-offer-2026-08-07.md`. IT IS A MINTING DEFECT, NOT A SCREEN-PREDICATE DEFECT, and that reverses the fix.** The screen was telling the truth: there genuinely was nothing recoverable under the key the box held, because the box minted that key itself over the top of a sealed package it already knew the hub was holding. **Three measurements carry it.** (1) **`WriteOffboxSecrets` (`offbox.go:411`) mints on ONE input — does the file exist.** It never reads `GetHubEscrowIdentityPresent()`, while its two neighbours in the same file, `OffsiteRecoveryOffer()` (`:1412`) and `needsOffsiteCredential()` (`:1377`), both do. **The same fact is available on three paths and used on two.** (2) **The flag was not merely available — it was the precondition of the chain that reached the minting.** The 5-minute retry job only logs when `RetryIfDeclared` fires, which requires the declaration, which requires that flag; the venue logged `credential retry: … (the box still declares a need; retrying)` at **02:48:03Z** and five times after — **thirty minutes and six ticks before the mint at 03:18:06Z**. (3) **The box KNEW and threw it away:** at **03:28:03Z**, thirty-five minutes before the customer looked, `EscrowAutoConfirmer.Reconcile` (`escrow_confirm.go:154`) computed the exact discriminator and logged `[WARN] [escrow-confirm] the hub's escrow blob does not cover the CURRENT repo password (hub hash 30ef574fe492… != local 9b4a9a9dcec7…) … staying pending`. **It is computed on every report cycle, never persisted, never surfaced.** **And the hub explicitly disclaims doing this** — `offsiteheal`'s package doc: *"it never runs, or asks for, an escrow ceremony… **credential automatic, key customer-present** — the ruling this session implements and must not quietly widen."* The repository key is minted in the seam between two sides that each honoured their contract, by a helper doing exactly what its doc comment says. **Fixing the predicate would paper over a box quietly making its own history unopenable.** **Recommended fix (not started, no code written): persist the discriminator the ACK already carries and add it as shape (c)** — needs no hub change and no new protocol field — **plus a `decided` latch**, because `ResetOrphanedRepo` clears `RepoState` without running a ceremony, so an `H`-mismatch discriminator alone would re-offer the screen forever to a customer who explicitly declined the old data. **Four operator decisions are stated and left unanswered in §"THE OPERATOR'S DECISION".** **✅ FIXED 2026-08-07 — controller v0.206.0 + hub v0.98.0, and the ruling above is what the fix follows.** **(1) It stops minting:** the guard is a CONJUNCTION (a package held AND no key present), so a first-time box mints exactly as before; the refusal is a HOLDING state, not a failure — the transport is still written so the recovery screen can bring the tier up the instant the key arrives (R-219), and returning an error instead would have left the hub re-staging a consumed credential for ever. New declared state `offsite.state=awaiting_recovery_key`, shown INERT to every existing hub reader from their code rather than assumed. **(2) The discriminator is persisted and drives the offer as shape (c).** §7.2 resolved deliberately: **a known difference offers however old the reading** (age is NOT gated on — gating would make a box offline from the hub silently stop offering, the very failure this removes), and **a hash never learned falls back to (a)/(b)**, because an empty hash is the hub positively saying its package seals no key rather than an unknown. **(3) Abandoning ends the question** — a 14-day countdown, visible and reversible, whose terminal step removes the set-aside store AND the sealed package together, after which shape (c) has nothing to compare and the offer falls silent **because the state is right, not because something remembers it once was not**. The two halves cannot be atomic across two machines, so it is a two-phase commit whose confirmation rides the SAME ACK that carries the request. **(4) The surface:** the full page appears **once per ENTRY into the offered state, not once ever** (an epoch — a box rebuilt months later is a new situation); three dismissal levers with three scopes, and **none removes the entry point on the backups page**. **Q7's trap does not survive:** while a recovery is outstanding „Helyreállítási kód létrehozása" is **unavailable**, not merely captioned. **§2.4 honoured:** the abandon confirmation no longer promises *„félretesszük — nem töröljük"* — it states the deletion date. **TWO REAL BUGS WERE CAUGHT BY TESTS RATHER THAN REVIEW, and both are recorded because the shape matters:** `OffboxAwaitingRecoveryKey` omitted `t.Enabled`, so a customer who had switched off-site OFF would have declared a holding state (caught by the EXISTING `TestOffsiteDeclare_DisabledTargetIsNotStranded`); and `recoveryInterrupts` returned early when the offer was false, so the FALLING edge was never recorded and the full page never came back — **the exact defect the epoch exists to fix, reintroduced inside the fix**. **Nine red-proofs, each with the mutation confirmed present in the file before its result was trusted**, including the ships-inert one (unwiring `RecordEscrowKeyHash`, which leaves everything compiling and every test passing while shape (c) reads an empty hash for ever). **NOTHING WAS DELETED ANYWHERE** — the terminal step has only ever run against injected fakes and an injected clock (§7.4). **Still open and NOT built by this:** R-242 (the release-to-golden gate) and R-245 (the automatic 30-day ending). | **FIXED 2026-08-07 — v0.206.0 / hub v0.98.0** | -| **R-242** | **A controller release that changes customer-visible behaviour is not delivered until a golden carries it — and nothing enforces that.** R-239 is the symptom; this is the mechanism, recorded **2026-08-07** and deliberately **NOT built** (the task that found it scoped it as a record-only item). **Two releases went out without a golden and the gap was invisible until a walk measured it from the customer's side**: v0.204.0 (R-237) and v0.205.0 (R-234) were written, tested, pushed, and CHANGELOG'd, and every one of those steps passed while a machine installed that night received neither. The register said CLOSED; the fleet said otherwise. **Nothing in the release path knows a golden exists.** The version bump, the image push, the CHANGELOG entry and the register closure are all repo-local; the manifest's `golden_version` is edited by a separate operator act, in a different repo, with no link back. **Proposed shapes, cheapest first — the choice is the operator's and is not taken here.** (a) **A release-path checklist step** — one line in the controller's end-of-session checklist: *a release that changes customer-visible behaviour is not finished until a golden carries it or a register row says why not.* Costs nothing, catches nothing mechanically. (b) **A gate in `repo_gates.py`** comparing the manifest's `golden_version` against the newest released controller and FAILING (or warning) past a tolerance of one minor. Mechanical, runs on every push, and would have fired the morning after v0.204.0. (c) **A hub-side checker** — the hub already knows every box's running controller version from `/hosts` and the vouched golden from the manifest; a periodic comparison against the newest published image would catch drift the repo cannot see, including a vouch that was made and then rolled back. **Earliest catch: (b).** It fires on the push that creates the gap, before any box is installed, and it needs no live fleet. **(c) catches strictly more but only after boxes exist.** (a) is worth doing regardless because it is free. **Not built. No gate was written this session.** | **READY — recorded, not built** — owner Viktor | +| **R-242** | **A controller release that changes customer-visible behaviour is not delivered until a golden carries it — and nothing enforces that.** R-239 is the symptom; this is the mechanism, recorded **2026-08-07** and deliberately **NOT built** (the task that found it scoped it as a record-only item). **Two releases went out without a golden and the gap was invisible until a walk measured it from the customer's side**: v0.204.0 (R-237) and v0.205.0 (R-234) were written, tested, pushed, and CHANGELOG'd, and every one of those steps passed while a machine installed that night received neither. The register said CLOSED; the fleet said otherwise. **Nothing in the release path knows a golden exists.** The version bump, the image push, the CHANGELOG entry and the register closure are all repo-local; the manifest's `golden_version` is edited by a separate operator act, in a different repo, with no link back. **Proposed shapes, cheapest first — the choice is the operator's and is not taken here.** (a) **A release-path checklist step** — one line in the controller's end-of-session checklist: *a release that changes customer-visible behaviour is not finished until a golden carries it or a register row says why not.* Costs nothing, catches nothing mechanically. (b) **A gate in `repo_gates.py`** comparing the manifest's `golden_version` against the newest released controller and FAILING (or warning) past a tolerance of one minor. Mechanical, runs on every push, and would have fired the morning after v0.204.0. (c) **A hub-side checker** — the hub already knows every box's running controller version from `/hosts` and the vouched golden from the manifest; a periodic comparison against the newest published image would catch drift the repo cannot see, including a vouch that was made and then rolled back. **Earliest catch: (b).** It fires on the push that creates the gap, before any box is installed, and it needs no live fleet. **(c) catches strictly more but only after boxes exist.** (a) is worth doing regardless because it is free. **Not built. No gate was written this session.** **⚠ IT RECURRED WITHIN A DAY, WHICH IS THE ARGUMENT FOR BUILDING IT.** Controller **v0.206.0** shipped the R-241 fixes on 2026-08-07 while the vouched golden still carried **0.205.0** — so a machine installed on the morning of 2026-08-08 would have received neither. Third occurrence of the shape in three days (R-111/R-115/R-120 are the older family). **✅ SHAPE (b) BUILT 2026-08-08 — `scripts/golden_currency_gate.py`, registered in `repo_gates.py` as gate 7.** **It was shown FAILING against that exact state before anything was baked**, which is its red-proof and the reason its own introducing push needed `--no-verify` (stated in the session report rather than worked around): `newest released controller : 0.206.0 / newest golden baked : 0.205.0 → CONVICTED`. **IT IS `--fast`, AND THAT FORCED ITS DESIGN:** both `.githooks/pre-push` AND CI run `repo_gates.py --fast`, so a non-fast gate would run in NEITHER — the R-29 census failure this runner exists to end. **THEREFORE IT CHECKS THE BAKE, NOT THE VOUCH**, because the vouched version lives only in the hub's `hub_settings` with no copy in git, and putting a copy there would create a second source of truth that can drift — a green gate over a false claim being the worst outcome available. **A bake without a vouch still passes: that half is NOT closed and stays on this row.** It also compares versions rather than behaviour, so a release changing nothing customer-visible trips it too — accepted deliberately, because judging that by hand is what failed three times and the cost of a false trip is one bake; a waiver belongs here, never in a habit of bypassing. | **PARTLY BUILT 2026-08-08 — the bake half is gated; the VOUCH half is not** | | **R-243** | **A box in the R-241 state silently stops backing up off-site, and NO ALARM OF ANY KIND FIRES.** Found by the R-241 spike (2026-08-07) as a by-product; **not part of the walk's finding and not previously filed.** The R-241 state is self-locking in a second, worse way than the recovery-journey dead end: `escrow_state` is stuck `pending` forever (the auto-confirm flips only on a hash match, and the hash cannot match a key the box minted itself), and `runOffboxBackup` returns at the escrow gate (`offbox.go:743`) before touching anything. **So off-site backups never run again — and the hub never notices.** All three signals that could catch it are excluded, each for its own individually-correct reason, verified in the hub this session: `offsite_stale` — `isStale` (`monitor/offsite.go:135`) returns false unless `EscrowState == "escrowed"`, and its own comment reads *"Pending/disabled = normal onboarding, never stale"*, so the box is classified as **still being set up, forever**; `offsite_delivery_stuck` — `monitor/offsite_delivery.go:91` skips the `applied` shape, and delivery genuinely IS applied (the credential was consumed and the target is in every report); `backup_failed` — never fires, because nothing fails: the run returns `nil` before it starts. **Three correct exclusions leaving one state unobserved.** This is the same class as the workspace `CLAUDE.md` "presence is not success" rule, one level up: **the absence of a failure is being read as the presence of a working tier.** **Partly subsumed by R-241's fix** — a box that recovers leaves this state — but **not for a box that does not**, and the alarm gap is what makes "does not" survivable indefinitely. **Not fixed; no code written.** **⚠ UPDATED 2026-08-07 (v0.206.0) — the STATE this row describes can no longer be entered, but the ALARM GAP is untouched and the row stays open.** R-241's mint guard means a box no longer mints a key over a sealed package, so it no longer arrives in the "escrow stuck pending against a self-minted key" state by itself. **What replaces it is a state that is VISIBLE rather than silent:** the box declares `offsite.state=awaiting_recovery_key` and the customer is offered the recovery screen. **But the hub still raises nothing for it**, and for the same three reasons: `isStale` needs `escrowed`, the delivery checker skips the `applied` shape, and `backup_failed` needs a run that never happens. **So a box whose customer never acts still stops backing up off-site with no operator signal** — the difference is that the customer can now see it and act, where before nobody could. **The remaining work is an operator-side signal for a box held in `awaiting_recovery_key` past some age**, and it is deliberately not bundled into R-241's fix. | **READY** — owner Viktor | | **R-244** | **The customer DELETE cascade leaves `app_log_issues` behind, and it is systematic across every venue ever torn down.** Found **2026-08-07** while verifying the `finalwalk` teardown with a **full census** (every table, every column) rather than a per-table query. After a cascade that logged `COMPLETE … full teardown`, **61 rows still matched `finalwalk`**. Four of the five sources are **deliberate and correct** — the cascade's own header states *"Provenance/events are NEVER wiped — audit outlives every tier"*: `events` 16, `notification_log` 14, `host_deletions` 1, `customer_resets` 1. **The fifth is a gap:** `app_log_issues` 29 rows, which the residue purge does not touch (its logged leg covers `reports`/`app_telemetry`/`app_log_tails`/`log_tail_requests`/`notif_prefs`/`selfbind_tokens`/`appliance_registrations` — not this table). **It is not a `finalwalk` quirk:** rows still reference **`c11` 40, `rewalk` 20, `part4` 24** — all three torn down 2026-08-06, whose ledger recorded *"0 occurrences"*. **That prior claim was measured with a narrower query and does not survive a full census; the correction is recorded rather than the measurement quietly redone.** **Why it was probably never written, established rather than assumed:** the table is a **fleet-wide aggregate** keyed on `app_name`+`fingerprint` with an `affected_customers` JSON list — of the 29 `finalwalk` rows, **12 reference only `finalwalk`** (orphans, safely deletable) and **17 are shared with LIVE customers** (`demo-felhom`, `peti-felhom`, …) and **must not be deleted, only de-referenced.** A naive `DELETE … WHERE customer LIKE` would destroy a live customer's issue history — which is very likely why the leg does not exist, and is the reason this is not a one-line fix. **Severity is LOW and stated plainly: no secret material is involved** — app name, fingerprint, message text, counts, timestamps. What survives is a deleted customer's *identifier* inside an aggregate row. **Proposed shape:** a residue leg that (a) removes the customer id from `affected_customers`/`context_customer`, and (b) deletes rows whose `affected_customers` becomes empty; plus a one-off sweep for the four already-torn-down venues. **The general lesson is the reusable part:** *a per-table absence query is not a census.* The teardown verification is now a full-schema sweep, and that is what found this. **Not fixed** — a cascade change needs its own red-proof and this session was scoped as a spike plus two operations. Evidence: `tests/teardown-finalwalk-2026-08-07.md`. | **READY** — owner Viktor | -| **R-245** | **Should a customer who never decides be auto-abandoned after 30 days? RECORDED, NOT BUILT — and the reasoning against it is recorded with it so the decision can be revisited properly.** **The operator's proposal (2026-08-07):** a box that has been offered recovery for 30 days without the customer deciding is auto-abandoned, entering the 14-day grace, so an undecided box does not sit for ever holding history nobody has claimed. **What was built instead:** the escalating reminders (1/3/7/14 days) and the operator levers `--abandon-extend` / `--abandon-stop`. **The reasoning, as settled with the operator the same day:** (1) **nobody is absent** — a box does not reinstall itself, so whoever rebuilt it was standing there and met the recovery question; the "customer away for months" case does not arise from this situation, because **a reinstall implies a person**. (2) **A customer who cannot find their code will get in touch**, which is the moment to extend or disarm by hand — so the automation would be firing at people we are already talking to, which is why the levers were the thing worth building. (3) **The cost is theirs**: the old history sits in the customer's own storage allowance, and if they are paying to keep something they have not decided about, that is their call and they feel it before we do. (4) **The real harm, if it comes, is QUOTA** — old history blocking new backups — and **that is a condition, not a calendar**. An automatic ending should trigger on the harm, with a dated warning, never on a date alone. **If this is ever built, build it that way.** **Do not implement it without a fresh ruling.** | **WAITING-ON-OPERATOR** — owner Viktor | +| **R-245** | **Should a customer who never decides be auto-abandoned after 30 days? RECORDED, NOT BUILT — and the reasoning against it is recorded with it so the decision can be revisited properly.** **The operator's proposal (2026-08-07):** a box that has been offered recovery for 30 days without the customer deciding is auto-abandoned, entering the 14-day grace, so an undecided box does not sit for ever holding history nobody has claimed. **What was built instead:** the escalating reminders (1/3/7/14 days) and the operator levers `--abandon-extend` / `--abandon-stop`. **The reasoning, as settled with the operator the same day:** (1) **nobody is absent** — a box does not reinstall itself, so whoever rebuilt it was standing there and met the recovery question; the "customer away for months" case does not arise from this situation, because **a reinstall implies a person**. (2) **A customer who cannot find their code will get in touch**, which is the moment to extend or disarm by hand — so the automation would be firing at people we are already talking to, which is why the levers were the thing worth building. (3) **The cost is theirs**: the old history sits in the customer's own storage allowance, and if they are paying to keep something they have not decided about, that is their call and they feel it before we do. (4) **The real harm, if it comes, is QUOTA** — old history blocking new backups — and **that is a condition, not a calendar**. An automatic ending should trigger on the harm, with a dated warning, never on a date alone. **If this is ever built, build it that way.** **✅ RE-FILED 2026-08-08 AS A DECISION TAKEN, not a question pending.** It sat in the operator's queue as `WAITING-ON-OPERATOR` for a day, and **nothing was actually pending** — the operator and the reviewer settled it on 2026-08-07: it is **not built**, the levers were built instead, and the whole reasoning above is the record of why. A settled decision parked in a queue is a queue nobody trusts, and an audit of every `WAITING-ON-OPERATOR` row the same day found this was the ONLY one — so the drift was caught while it was still a single row. **THE CONDITION THAT REOPENS IT, which the reasoning already names: QUOTA — old set-aside history blocking new backups.** Not a calendar. If a customer's retained history ever refuses a new backup, revisit this with a dated warning that triggers on the refusal; until then it stays decided. | **DECIDED 2026-08-07 — not built; reopens on quota** | + +| **R-246** | **A leftover staleness flag has silently disabled the new recovery discriminator on `demo-hp` since 2026-08-04, and the flag is WRONG.** Found by a read-only spike, 2026-08-08. **Q1 — traced to an act, to the second:** at `2026-08-04 20:15:49` the hub emitted `offsite_reissued` and `escrow_stale` in the same second — an operator **Re-issue** pressed during the R-201 drill, three minutes after `escrow_blob_served` at 20:12:40/20:12:54. That was `offsite.ReissueCredentials`'s **precautionary** `MarkEscrowStale` call, which **hub v0.95.0 REMOVED the very next day** (R-196 / R-204 item 2) precisely because it marked healthy escrows stale. **Q2 — the flag is wrong, measured on both sides:** the hub's blob seals `restic_pw_sha256 = 8a9e33aa4da6769c…d080a`, and the key the box is actually using hashes to **the identical value**. The blob covers the key. **Q3 — nothing clears it by itself:** the ONLY writer of `stale_at = NULL` is `SaveHostEscrow`'s `ON CONFLICT` — i.e. a fresh escrow ceremony, **which is the one act that would supersede the good blob**. So the only exit from a false alarm is the destructive act the false alarm recommends. **Q5 — the blast radius, enumerated rather than assumed:** (1) `GetEscrowStatusForCustomer` withholds `restic_pw_sha256` from the ACK; (2) a **pending** box can never auto-confirm, so (3) **every off-site run is refused indefinitely** — neither bites `demo-hp`, which was already `escrowed` and is backing up healthily (12 snapshots, last success 2026-08-07T02:15:35Z); (4) the customer is told to create a new code; (5) **NEW — v0.206.0's shape (c) is inert**, because the box records an empty hub hash and falls back to (a)/(b), so the recovery screen would stay silent even if recovery were needed. **Q6 — a fresh box CANNOT reach this state:** `MarkEscrowStale` has **no production caller anywhere in the tree** (census: only its own definition, two comments and two test references). **The next walk cannot meet it.** **NOT CLEARED, deliberately** — Q2's answer says the fix is to clear this instance, but whether to also stop the column being settable at all is a separate ruling, and the spike was scoped read-only. **What it needs:** an operator decision to clear `stale_at` for `demo-hp-bb76ea` (a one-row UPDATE), and a ruling on whether the column keeps a live setter. | **READY** — owner Viktor | +| **R-247** | **The box is being told something false, in its own words, and it recommends the destructive act.** `demo-hp`, live, on controller v0.206.0: *"STALE escrow: the hub's current blob carries NO password hash (hash-less supersession) — the stored recovery bundle does not cover the offsite password; create a new recovery code"*. **Every clause is false.** The hub HAS the hash and is *withholding* it (R-246); there was no supersession (`host_escrow_superseded` has no row for this host); and the bundle **does** cover the password — the hashes match exactly. It raises `EscrowStale`, which renders the customer-facing card *„A letétben lévő helyreállítási csomag nem fedi a jelenlegi távoli mentési jelszót. Hozzon létre új helyreállítási kódot."* **THE CAUSE IS A FACT ON THE WIRE THAT THE BOX THROWS AWAY.** The hub already sends `escrow_stale` in the ACK (`json:"escrow_stale,omitempty"`), and the controller's `report.EscrowStatus` **has no matching field**, so `encoding/json` drops it silently. The box therefore cannot distinguish *withheld because flagged stale* from *genuinely hash-less*, and guesses the latter. **This is R-241's shape for the third time: the answer is available, and it is discarded at the boundary.** **The fix is small and is NOT made here** (§0 forbids a controller change this session): add the field, and say the true thing — or say nothing, since on an ESCROWED box with matching hashes there is nothing wrong to report. | **READY** — owner Viktor | +| **R-248** | **A flag that changes behaviour is visible to nobody who would look for it.** Q4 of the 2026-08-08 spike, answered plainly. **The customer** sees only a derived card stating a false reason (R-247). **The box** cannot see it at all (R-247's dropped field). **The operator** can see it on exactly ONE page — the **PBS-DR** view (`hub/internal/web/pbsdr.go:487`, `v.EscrowStale = escrow.StaleAt != ""`) — which is the wrong tier for this symptom: an operator investigating an OFF-SITE problem has no reason to open a PBS-DR page. **No alert, no report field, no off-site surface.** The one-shot `escrow_stale` event fired on 2026-08-04 and **was never notified** (a full census of `notification_log` for that customer that day returns 8 rows, none of them this one); it has fired twice ever, both on 4 August. **So the practical answer is: only a database read.** That is a finding in its own right — a flag that silently changes behaviour and cannot be seen is the shape this fortnight has been about (R-241's discarded comparison, R-228's unread field, R-243's unobserved state). **What it needs:** surface `stale_at` on the off-site operator surface and in the host report, or stop using a field nobody can observe to change what a customer is told. | **READY** — owner Viktor | **Recorded against existing rows by Phase 2:**