Files
felhom.eu/REPORT.md
T
admin c04f933d0b
gates / gates (push) Successful in 13s
The census answers no, three receipts found, and the prune was on file all along
CENSUS (read-only, hub store, tester's machine not contacted): no machine that is
not ours can be in the state that cost demo-felhom its history. The hub holds
escrow for three hosts; both demo boxes lost their pre-fix key in the same four
hours on 2026-08-04; peti-felhom and david have no host row and no escrow at all.
A control ran FIRST and had to pass -- the query returned "present (572 bytes)"
for a host known to have material and "absent (NULL)" for one known not to.
Corrected my own instrument on the way: a date-only comparison mislabelled both
losses as after the fix, so the in-force moment is now pinned from the hub's first
post-fix escrow row (11:11:37Z), which independently agrees with the register.

PART 1 ESTABLISHED. The prune is recorded inside R-267 -- the row about the
Configuration page being slow -- because pruning artifacts is what made that page
fast. Arithmetic checks (23+7=30, plus three versions that only surfaced after the
first thirty moved them onto page one = 33) and the PAGINATED listing shows both
generics at exactly ten. R-291's blocking condition is released: the operator was
being asked to establish something already written down. And my counter-argument
yesterday was wrong in exactly the way R-267 warns about -- "containers hold 19"
came from an unpaginated query; paginated they hold 270 and 169.

RECEIPTS: three restored (drives.enrol, backup.tier1, fail.lost-recovery-code),
each citing the document that walked it; the map already read PROVEN-LIVE for all
three, so this follows the map rather than raising a status in the view. NINE
HONEST GREYS. fault.selfheal's best hit argues against it -- an incident recording
self-heal's absence through a 1h15m outage.

THE DECAY RULE FIRED FOR THE FIRST TIME. backup.restore-proof has a receipt from
28 July and is superseded anyway: demo-hp's restore-test failed 5 August and the
box has since been rebuilt. A claim about a continuing behaviour cannot rest on an
old observation. The capability map still reads PROVEN-LIVE and is now the thing
out of step -- recorded, not silently rewritten.

PART 4 specified, not implemented. The orphan card promises restorability the box
rendering it cannot evaluate: the discriminator is on the hub and no wire field
carries it. A conditional promise the system cannot evaluate is the same defect as
an unconditional false one, so the copy stops promising, says what happens, and
names a route. Ships with the next controller change so one bake covers both.
2026-08-10 11:15:37 +02:00

9.4 KiB
Raw Blame History

REPORT — the census, the receipts, and a promise we should stop making (2026-08-10)

Documentation and one read-only census. No code that runs on a customer's machine, no deploy, no bake. Both demo boxes are online and were not changed. The tester's machine was not contacted — only its hub-side record was read, and it has none.


1. Part 2's answer, first

No. A machine that is not ours cannot be in the state that cost demo-felhom its history — not today, and not for anything enrolled from now on.

The hub holds escrow for three hosts. Both demo boxes lost their pre-fix key in the same four-hour window on 2026-08-04; that is the entire affected population and it is entirely ours. peti-felhom has no host row and no escrow at all, and neither does david. Anything enrolled from here escrows under the fixed code, which has been in force since 2026-08-04.

2. The census, and the control that came first

Control (run before the census, and it had to pass or the census was worthless):

demo-felhom-8363b5  host_escrow             -> present (572 bytes)   want present   OK
demo-felhom-8363b5  host_escrow_superseded  -> absent (NULL)         want absent    OK
control PASSED: the query distinguishes both states on known cases.
host current superseded material superseded_at verdict
demo-felhom-8363b5 572 B id=4 ABSENT 2026-08-04 07:20:08 old backups lost (before the fix)
demo-hp-bb76ea 572 B id=3 ABSENT 2026-08-04 07:15:36 old backups lost (before the fix)
drill-r50-0a4f9a none 0 no supersession has happened
peti-felhom no host record, no escrow
david no host record, no escrow

An orphan check found no escrow row pointing at a host the hub does not know.

One correction I made to my own instrument. The first run labelled both losses "superseded AFTER the fix — unexpected", because a date-only comparison puts 2026-08-04 07:20:08 after 2026-08-04. The in-force moment is pinned instead from the hub's own first post-fix escrow row — 11:11:37Z — which independently agrees with the register's "four hours before v0.93.0 fixed the retention". Both losses are then correctly before the fix: explained, not anomalous.

What this does NOT claim: that a retained blob has ever been unwrapped on a superseded row. Retention is proven; recovery from a superseded row remains unexercised.

3. Part 1 — the deleter is ESTABLISHED, and it was on file all along

The locator was correct: the record is inside the R-267 row"Pruned to the newest 10 per package on the operator's rule, with the live-vouched golden/agent/floor asserted into the KEEP set before a single DELETE was issued; 33 deletions, all HTTP 204."

Every corroboration checked and every one holds:

  • The arithmetic: 23 agent + 7 golden = 30, plus three older agent versions (0.81.0/0.80.0/0.79.0) that "only became visible after the first 30 deletions moved them onto page one" = 33.
  • The live PAGINATED listing — 14 pages, 653 package-versions — felhom-agent generic at exactly 10, felhom-golden generic at exactly 10. That is what a newest-10 prune leaves.

The midnight-cleanup candidate is retired, and R-291's blocking condition is released: the operator was being asked to establish something already written down.

My counter-argument yesterday was wrong, in precisely the way R-267 warns about. I argued against the prune because "container packages hold 19 each". That came from an unpaginated query the API caps at 50/page. Paginated, they hold 270 and 169 — they were never in the prune. R-267 records the identical trap one paragraph above the sentence I could not find: "An unpaginated listing is not evidence of a total — this repo's own rule, walked into while measuring."

4. Part 3 — the twelve

Three restored. Nine honest greys. The map already read PROVEN-LIVE for the three, so the dataset was behind it — restoring follows the map rather than raising a status in the view.

claim outcome
drives.enrol RESTOREDaudits/SPIKE-raw-drive-enroll-2026-06-15.md: a live drive walked scan → format → mount → PVE storage → one-click enrol, with the resulting storage.cfg entry and mount unit recorded
backup.tier1 RESTOREDaudits/CAMPAIGN-8-backup-restore-2026-07-27.md: adversarial, destructive, unattended, both boxes + ep0; A2 proven end to end
fail.lost-recovery-code RESTOREDaudits/REHEARSAL-byo-reinstall-2026-08-09.md, and proven the hard way on 2026-08-10
backup.restore-proof grey — THE DECAY RULE FIRED (below)
install.installer-by-tag grey — hits are ISO spikes; nothing walks a tag rollback
use.lifecycle grey — two passing mentions, no walk
drives.migrate grey — one or two mentions only
backup.whole-machine grey — diagnostics and phase findings, no walk of the claim
fault.selfheal grey — the best hit argues the other way: INCIDENT-guest-dhclient-killed-2026-07-20.md documents self-heal's absence through a 1 h 15 m outage
fault.operator-email grey — source-verified, delivery never observed
fail.drive-filling grey — weak hits only
fail.hub-down grey — Campaign 11's hub-unreachable work produced a finding (R-224), not a pass

5. Did the "code moved under the proof" rule fire? Yes — for the first time.

backup.restore-proof has a receipt: architecture/_recovery-inventory-2026-07-28.md carries live journal lines for scheduled restore-tests on both boxes and both tiers. It is superseded anyway — demo-hp logged restore_test_failed on 2026-08-05, and the box has since been wiped and reinstalled. The claim is about a continuing scheduled behaviour, so a 2026-07-28 observation cannot carry it.

So the rule can fire, and now has. Two nights ago it fired zero times out of twelve because every downgrade came from missing evidence, not decayed evidence — there was nothing for it to bite on. A consequence to note: the capability map still reads PROVEN-LIVE (2026-08-03) for that row, so the map is now the thing out of step, and it is the source. Recorded rather than silently rewritten.

6. Part 4 — the promise, and why the wording is what it is

Surfaces:

file:line
controller/internal/web/templates/backups_remote.html:101 the false promise
controller/internal/web/templates/backups_remote.html:98 the orphan explanation — accurate, keep
controller/internal/web/templates/layout.html:143 the 14-day abandon countdown — accurate, keep
controller/internal/settings/settings.go:337-338 OrphanedRenamedTo schema comment — accurate, keep

(The capability map was searched and makes no such claim — nothing to correct there.)

The three cases: set aside before the fix → not recoverable by construction (the restic password lives only in the identity bundle, escrow/identity.go:39 read by escrow/recover.go:91, and those rows are NULL); set aside after → recoverable in principle, never demonstrated; today's population is entirely the first case.

The deciding fact: neither the box nor the customer can tell which case they are in. The hub holds the discriminator (host_escrow_superseded.identity_blob); the box caches only HubEscrowIdentityPresent, which is about the current escrow; and no field on the report or ACK wire carries superseded-blob retention. The box renders the card. A conditional promise the renderer cannot evaluate is the same defect as an unconditional false one — so the specified copy stops promising, says plainly what happens, explains why it cannot promise, and names a route (write to us).

Specification: documentation/design/SPEC-orphan-card-copy-2026-08-10.md — copy, surfaces, and the render tests that should pin it, including a regression guard that the string visszaállítható lehet never returns. Not implemented, on purpose: it lands in the controller, and a controller release is undelivered until a golden carries it (R-242). It should ship with the next controller change so one bake and one approval cover both, and the spec says so.

7. Register

Ceiling R-292 → R-294. Opened R-293 (the census) and R-294 (the promise + spec). R-287 turned to ESTABLISHED with my unpaginated-count error withdrawn. R-291's blocking condition released. R-288 gained a second measured cost, and it is a different failure mode from the first: prose rows make claims ambiguous; rows-about-other-things make facts unfindable.

8. Observations — noticed, not acted on

  • The capability map is now out of step in two directions — behind the dataset for three claims it already called PROVEN-LIVE, and ahead of it for backup.restore-proof. Both point at R-288.
  • fault.selfheal's only real document argues against it. Worth someone deciding whether the capability is real and unwalked, or overstated.
  • Retention is proven; recovery from a superseded row is not. The census proves blobs are now kept; nobody has ever unwrapped one. That is the next thing worth a drill, and it needs no customer.
  • The page now shows the three restored claims as "moved twice" rather than once, so an unsettled status reads as unsettled.