The census answers no, three receipts found, and the prune was on file all along
gates / gates (push) Successful in 13s

CENSUS (read-only, hub store, tester's machine not contacted): no machine that is
not ours can be in the state that cost demo-felhom its history. The hub holds
escrow for three hosts; both demo boxes lost their pre-fix key in the same four
hours on 2026-08-04; peti-felhom and david have no host row and no escrow at all.
A control ran FIRST and had to pass -- the query returned "present (572 bytes)"
for a host known to have material and "absent (NULL)" for one known not to.
Corrected my own instrument on the way: a date-only comparison mislabelled both
losses as after the fix, so the in-force moment is now pinned from the hub's first
post-fix escrow row (11:11:37Z), which independently agrees with the register.

PART 1 ESTABLISHED. The prune is recorded inside R-267 -- the row about the
Configuration page being slow -- because pruning artifacts is what made that page
fast. Arithmetic checks (23+7=30, plus three versions that only surfaced after the
first thirty moved them onto page one = 33) and the PAGINATED listing shows both
generics at exactly ten. R-291's blocking condition is released: the operator was
being asked to establish something already written down. And my counter-argument
yesterday was wrong in exactly the way R-267 warns about -- "containers hold 19"
came from an unpaginated query; paginated they hold 270 and 169.

RECEIPTS: three restored (drives.enrol, backup.tier1, fail.lost-recovery-code),
each citing the document that walked it; the map already read PROVEN-LIVE for all
three, so this follows the map rather than raising a status in the view. NINE
HONEST GREYS. fault.selfheal's best hit argues against it -- an incident recording
self-heal's absence through a 1h15m outage.

THE DECAY RULE FIRED FOR THE FIRST TIME. backup.restore-proof has a receipt from
28 July and is superseded anyway: demo-hp's restore-test failed 5 August and the
box has since been rebuilt. A claim about a continuing behaviour cannot rest on an
old observation. The capability map still reads PROVEN-LIVE and is now the thing
out of step -- recorded, not silently rewritten.

PART 4 specified, not implemented. The orphan card promises restorability the box
rendering it cannot evaluate: the discriminator is on the hub and no wire field
carries it. A conditional promise the system cannot evaluate is the same defect as
an unconditional false one, so the copy stops promising, says what happens, and
names a route. Ships with the next controller change so one bake covers both.
This commit is contained in:
2026-08-10 11:15:37 +02:00
parent 67eced8fbf
commit c04f933d0b
8 changed files with 344 additions and 240 deletions
+56 -87
View File
@@ -27,6 +27,52 @@ So I took your stated fallback: the old store was **moved aside, not deleted** (
**One thing that needs your judgement, not mine.** The card that offered this told the customer their set-aside backups *may be restorable later with their recovery code*. For these ones that is simply untrue, and it is said to precisely the people who have just lost their history. *(R-202 — now evidenced.)*
## Can anyone else lose their history the way demo-felhom did? No.
**One read of the hub's own records, no machine touched.** The hub holds backup keys for exactly
**three** machines. Both demo boxes lost their old key in the same four-hour window on 4 August,
before the retention fix was in force — that is the whole population of the problem, and it is
entirely ours. **The tester's machine has no record at all**, so it cannot be affected; and anything
enrolled from now on is covered, because the fix has been in force since 4 August.
I ran a control before trusting the query: it had to say *material present* for a machine known to
have it and *absent* for one known not to. It did both.
## Three green dots came back, nine stayed grey, and one rule finally fired
**The nine greys are the honest number.** For those, no document anywhere walks the claim, and saying
so is more useful than a dot nobody can defend.
- **Back to green**, each citing the document that walked it: the drive wizard (a live drive taken
through scan → format → mount → enrol), the on-box app backups (an overnight destructive campaign
across both machines), and the lost-recovery-code case — which we then proved the hard way this
morning.
- **One claim stayed grey for a new reason, and it is the interesting one.** The unattended
restore-proof *does* have a receipt from 28 July — but demo-hp's restore-test failed on 5 August and
the machine has since been wiped and rebuilt. It is a claim about something that keeps happening, so
an old observation cannot carry it. **This is the first time that rule has fired**; two nights ago it
fired zero times out of twelve.
## The prune mystery is solved, and the answer was written down all along
Who deleted the old versions: **you did, on 4 August evening, on your own rule** — 33 deletions, keep
set asserted first, every one a clean 204. **It was recorded inside the row about the Configuration
page being slow**, because pruning artifacts is what made that page fast. Two sessions failed to find
it. **You are no longer blocked** on establishing something that was already on file.
I also got a number wrong yesterday and it is corrected: I said the container packages held nineteen
versions and used that to argue against the prune. Counted properly — with pages — they hold 270 and
169, and the two that *were* pruned sit at exactly ten each.
## The sentence we should stop saying
When a machine's off-site history is set aside, the card tells the customer it *may be restorable
later with their recovery code*. **The machine showing that card cannot know whether it is true**
the fact lives on the hub and is not sent to the box. For anything set aside before 4 August it is
simply false. **The replacement wording is written and waiting**
(`documentation/design/SPEC-orphan-card-copy-2026-08-10.md`); it ships with the next controller
release so one image bake and one approval cover it, rather than costing you two of each.
## What works
A blank machine boots the Felhom disc, installs itself unattended, and is claimed by the customer, who
@@ -40,74 +86,18 @@ destroyed on purpose and its files came back byte for byte identical — four ti
their recovery code got everything back with **no command line inside the machine at any point**,
in 72 seconds. The two rough edges that walk found are also gone. *(R-201, R-252, R-253 — closed.)*
## The guard is in. An approval that cannot be installed is now refused.
## Shipped 2026-08-09 — the guards, and the rehearsal that earned them
**Yesterday morning every install in existence failed for hours, and nothing would have stopped it
happening again. Now something does.** When you press Save on the Day-0 artifacts, the hub checks —
before it writes — that each version you are vouching actually has its git label **and** that its file
can actually be downloaded. If either is missing it refuses and tells you which, for which version,
and the one command that fixes it. **Hub v0.102.0 is live.**
**An approval that cannot be installed is now refused** at the moment you press Save: the hub checks
the version's git label and that its file downloads, and refuses with a message naming the fix. A
second machine catches it a step earlier in the agent repo. Both were owed after every install in
existence failed for hours on 2026-08-09. **Hub v0.102.0 is live.**
Three details worth your knowing:
- **It checks the exact file the installer fetches first** — the one whose absence broke Friday — not
some other file that happens to exist. A test pins that, because probing the wrong file is precisely
how the failure stayed invisible.
- **"Could not check" also refuses**, with a different message. Saving with a warning would read as a
success, and we have the scars. **There is no override**: the registry is on your own server, so if
it is unreachable the approval can wait.
- **A second machine now catches it one step earlier** — the agent repo refuses to consider a release
complete unless its label and its package both exist.
## The red repository is green, and the two rules now share one number
The tidy-up keeps the newest ten versions; the check demanded every version ever labelled still be
downloadable. Both are sensible and together impossible, so the red would have returned on your next
publish. They now read the same number from one file. **What the check no longer covers, plainly: a
version older than the ten is no longer asserted downloadable** — its label and its config files still
are — and it prints which ones it dropped on every run so this cannot go quiet.
**One thing I could not establish, and I am not guessing.** Who actually deleted the old versions is
**still unknown**. Gitea keeps no deletion trail: no cleanup rule is configured, the version table has
no deleted-marker, the activity feed shows no package operation, and the server log no longer reaches
back that far. I also **withdrew my own claim from yesterday** that the logs showed no deletion — the
logs did not cover the window, so they never said anything.
## The rehearsal finished. The data came back byte for byte; the journey did not.
**We wiped a working demo machine and put it back. All four test files returned identical — including
the two with Hungarian accents, checked as raw bytes, not as text on screen.** The unlock took 21
seconds and the restore 13. **But it only finished because I could open a terminal twice.** A
household would have stopped, twice, and the second time the screen would have told them it was easy.
**The two walls, both fixed-or-fixable, neither about the data:**
- **The install died four steps in** — the agent version you approved had been published as a download
but never given its version label, and the installer looks it up by that label. **Now unblocked**
I pushed the label after checking the published file matched what you vouched. *(R-273 — closed. The
two guards that would stop it recurring are still owed.)*
- **A reinstalled machine cannot re-attach its own data drive.** Every route is a dead end, and the
restore page cheerfully says „**Ez két kattintás**" while pointing at an empty list. The drive is
fine and the machine can see it — it just is not offered, because the same drive is also the backup
target. I got past it by typing an internal path no customer could know. **This is the one to fix
before the tester's visit.** *(R-280)*
**Also broken, found on the way:**
- **Taking Felhom off a machine leaves the one thing that stops it going back on.** We install a small
network service at setup; removing Felhom restarts it without its settings, it seizes the port the
next install needs, and the next install then refuses — appearing to blame the owner's network.
*(R-272)*
- **A machine we removed keeps its private line to us open.** *(R-276)*
- **The hub said nothing at all** while a machine was wiped, rebuilt, re-claimed and had its sealed
backups opened. No false alarm — but also no word, and the alarm that exists for "someone is opening
this customer's backups" stayed silent through a real one. *(R-281)*
- **A rebuilt machine may still come back on software from last week** — narrower than I first wrote:
the resumed install fetched the right version, but a fresh one takes whatever copy is newest on the
disk without checking it against what you approved. *(R-274)*
- **demo-felhom has not had an off-site backup in six days** and is waiting for a recovery code nobody
has entered. demo-hp, rebuilt the same day, recovered by itself. *(R-278)*
- **One code, three different names**, and the email points at a page the machine is not showing —
this cost us a wasted code today. *(R-282, R-283)*
**The reinstall rehearsal:** a demo machine was wiped and put back. All four test files returned
**byte for byte**, accented Hungarian filenames included, checked as raw bytes. But it only finished
because a terminal was available twice — the install died on a missing version label *(R-273, now
guarded)* and **a reinstalled machine still cannot re-attach its own data drive** *(R-280 — the one to
fix before the tester's visit)*. Detail: `documentation/audits/REHEARSAL-byo-reinstall-2026-08-09.md`.
## What's broken
@@ -121,27 +111,6 @@ household would have stopped, twice, and the second time the screen would have t
accumulates. *(R-244)*
- **Putting restored files back where they belong is still manual.** *(R-213)*
## The rest of what the rehearsal found
Sixteen findings in one afternoon, **none of them visible from reading the code** — three sessions of
review had not seen any.
- Removing Felhom leaves five files holding old keys *(R-275)*, and rotating a leaked key does not
revoke the old one until the service restarts *(R-269)* — the written recipe for it is a step short
*(R-270)*, and the alarm it raises can never be closed because the fix it recommends is what
silences the all-clear *(R-271)*.
- **It caught me being wrong twice, and that matters more than the count.** I told you the fleet's
off-site backups were down; demo-hp had eighteen snapshots and I had read three misleading screens
instead of asking the machine *(R-277)*. And I raised a leftover permissions file as a security
hole, then tested it and refuted myself — it is inert.
**What worked, and should not be lost in the count:** the machine came up on its own at the approved
version; the setup page appeared unprompted, in Hungarian, naming the customer; **the recovery screen
appeared without being looked for** and said plainly that unlocking changes nothing; the restore told
the truth about putting files in a checking folder rather than back in place; and no false alarm fired.
Full account: `documentation/audits/REHEARSAL-byo-reinstall-2026-08-09.md`.
## Three rulings, written down so they stop living in a conversation
- **The managed-update floor.** It is deliberately parked, and the trigger to raise it is **the first