The census answers no, three receipts found, and the prune was on file all along
gates / gates (push) Successful in 13s

CENSUS (read-only, hub store, tester's machine not contacted): no machine that is
not ours can be in the state that cost demo-felhom its history. The hub holds
escrow for three hosts; both demo boxes lost their pre-fix key in the same four
hours on 2026-08-04; peti-felhom and david have no host row and no escrow at all.
A control ran FIRST and had to pass -- the query returned "present (572 bytes)"
for a host known to have material and "absent (NULL)" for one known not to.
Corrected my own instrument on the way: a date-only comparison mislabelled both
losses as after the fix, so the in-force moment is now pinned from the hub's first
post-fix escrow row (11:11:37Z), which independently agrees with the register.

PART 1 ESTABLISHED. The prune is recorded inside R-267 -- the row about the
Configuration page being slow -- because pruning artifacts is what made that page
fast. Arithmetic checks (23+7=30, plus three versions that only surfaced after the
first thirty moved them onto page one = 33) and the PAGINATED listing shows both
generics at exactly ten. R-291's blocking condition is released: the operator was
being asked to establish something already written down. And my counter-argument
yesterday was wrong in exactly the way R-267 warns about -- "containers hold 19"
came from an unpaginated query; paginated they hold 270 and 169.

RECEIPTS: three restored (drives.enrol, backup.tier1, fail.lost-recovery-code),
each citing the document that walked it; the map already read PROVEN-LIVE for all
three, so this follows the map rather than raising a status in the view. NINE
HONEST GREYS. fault.selfheal's best hit argues against it -- an incident recording
self-heal's absence through a 1h15m outage.

THE DECAY RULE FIRED FOR THE FIRST TIME. backup.restore-proof has a receipt from
28 July and is superseded anyway: demo-hp's restore-test failed 5 August and the
box has since been rebuilt. A claim about a continuing behaviour cannot rest on an
old observation. The capability map still reads PROVEN-LIVE and is now the thing
out of step -- recorded, not silently rewritten.

PART 4 specified, not implemented. The orphan card promises restorability the box
rendering it cannot evaluate: the discriminator is on the hub and no wire field
carries it. A conditional promise the system cannot evaluate is the same defect as
an unconditional false one, so the copy stops promising, says what happens, and
names a route. Ships with the next controller change so one bake covers both.
This commit is contained in:
2026-08-10 11:15:37 +02:00
parent 67eced8fbf
commit c04f933d0b
8 changed files with 344 additions and 240 deletions
+24
View File
@@ -15,6 +15,30 @@
> would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`** > would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`**
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below. > and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
## A record that lives inside a row about something else has not been recorded
**Earned 2026-08-10, at a measured cost of two sessions.** Thirty-three package deletions — an
operator-ruled prune, with the keep set asserted first and every delete returning 204 — were written
down inside **R-267, the row about the Configuration page being slow**, because pruning artifacts is
what made that page fast. It was a perfectly good record in a place nobody would look.
**What it cost.** One session searched the register and reported *"no register row records a package
prune"*. A second exhausted Gitea's router logs, its activity feed and its schema, then concluded the
deleter *"may be unestablishable from this side"*. Both were wrong, and the answer was in
`OPEN-ITEMS.md` throughout. Worse, the second session argued against the prune using an **unpaginated**
package count — the exact trap that same R-267 row documents one paragraph above the sentence it could
not find.
**Findability is part of recording, not a nicety.** A fact filed under the story of how it was
discovered is filed under the wrong thing. Concretely:
- **An execution record gets its own row**, even when the execution was incidental to something else.
A row may cross-reference; it may not be the only home.
- **A row about a fix is not the home for the acts the fix required.**
- This is the second measured cost of the register's illegibility in two days, and it is a *different*
failure from the first: prose rows make claims ambiguous, rows-about-other-things make facts
unfindable. Both are in R-288.
## A check and the policy it enforces must read the same number from the same place ## A check and the policy it enforces must read the same number from the same place
**Earned 2026-08-09.** The container registry retains N versions per package; `check-published-versions.py` **Earned 2026-08-09.** The container registry retains N versions per package; `check-published-versions.py`
+129 -131
View File
@@ -1,152 +1,150 @@
# REPORT — the guards, and one thing I could not establish (2026-08-09 evening) # REPORT — the census, the receipts, and a promise we should stop making (2026-08-10)
No machine touched. Both demo boxes stayed offline and in transit; nothing here needed them. Documentation and one read-only census. **No code that runs on a customer's machine, no deploy, no
bake.** Both demo boxes are online and were not changed. **The tester's machine was not contacted**
**Parts 1, 2 and 3 are done. Part 4 and Part 5 are NOT done** — see §6 and §8. Part 5 was marked only its hub-side record was read, and it has none.
droppable and was dropped; **Part 4 was not marked droppable and I ran out of session before it**,
which is a shortfall rather than a decision.
--- ---
## 1. Part 1.1 — the deleter: **NOT ESTABLISHED** ## 1. Part 2's answer, first
**This prompt's account is not corroborated by any source I can reach**, and the prompt itself said > **No. A machine that is not ours cannot be in the state that cost demo-felhom its history — not
not to take its word for it. > today, and not for anything enrolled from now on.**
| claim in the prompt | what I found | The hub holds escrow for **three** hosts. Both demo boxes lost their pre-fix key in the same four-hour
window on 2026-08-04; that is the entire affected population and it is entirely ours. `peti-felhom`
has **no host row and no escrow at all**, and neither does `david`. Anything enrolled from here escrows
under the fixed code, which has been in force since 2026-08-04.
## 2. The census, and the control that came first
**Control (run before the census, and it had to pass or the census was worthless):**
```
demo-felhom-8363b5 host_escrow -> present (572 bytes) want present OK
demo-felhom-8363b5 host_escrow_superseded -> absent (NULL) want absent OK
control PASSED: the query distinguishes both states on known cases.
```
| host | current | superseded | material | superseded_at | verdict |
|---|---|---|---|---|---|
| `demo-felhom-8363b5` | 572 B | id=4 | **ABSENT** | 2026-08-04 07:20:08 | old backups lost (before the fix) |
| `demo-hp-bb76ea` | 572 B | id=3 | **ABSENT** | 2026-08-04 07:15:36 | old backups lost (before the fix) |
| `drill-r50-0a4f9a` | none | 0 | | | no supersession has happened |
| `peti-felhom` | — | — | — | — | **no host record, no escrow** |
| `david` | — | — | — | — | **no host record, no escrow** |
An orphan check found no escrow row pointing at a host the hub does not know.
**One correction I made to my own instrument.** The first run labelled both losses *"superseded AFTER
the fix — unexpected"*, because a date-only comparison puts `2026-08-04 07:20:08` after `2026-08-04`.
The in-force moment is pinned instead from the hub's own first post-fix escrow row — **11:11:37Z**
which independently agrees with the register's *"four hours before v0.93.0 fixed the retention"*. Both
losses are then correctly *before* the fix: explained, not anomalous.
**What this does NOT claim:** that a *retained* blob has ever been unwrapped on a superseded row.
Retention is proven; recovery **from** a superseded row remains unexercised.
## 3. Part 1 — the deleter is ESTABLISHED, and it was on file all along
The locator was correct: the record is **inside the R-267 row** — *"Pruned to the newest 10 per package
on the operator's rule, with the live-vouched golden/agent/floor asserted into the KEEP set before a
single DELETE was issued; 33 deletions, all HTTP 204."*
Every corroboration checked and every one holds:
- **The arithmetic:** 23 agent + 7 golden = 30, plus three older agent versions (0.81.0/0.80.0/0.79.0)
that *"only became visible after the first 30 deletions moved them onto page one"* = **33**.
- **The live PAGINATED listing** — 14 pages, 653 package-versions — `felhom-agent` generic at **exactly
10**, `felhom-golden` generic at **exactly 10**. That is what a newest-10 prune leaves.
The midnight-cleanup candidate is retired, and **R-291's blocking condition is released**: the operator
was being asked to establish something already written down.
**My counter-argument yesterday was wrong, in precisely the way R-267 warns about.** I argued against
the prune because *"container packages hold 19 each"*. That came from an **unpaginated** query the API
caps at 50/page. Paginated, they hold **270** and **169** — they were never in the prune. R-267 records
the identical trap one paragraph above the sentence I could not find: *"An unpaginated listing is not
evidence of a total — this repo's own rule, walked into while measuring."*
## 4. Part 3 — the twelve
**Three restored. Nine honest greys.** The map already read PROVEN-LIVE for the three, so the dataset
was *behind* it — restoring follows the map rather than raising a status in the view.
| claim | outcome |
|---|---| |---|---|
| "the register records … pruned to the newest ten … 33 deletions, all HTTP 204" | **No such row exists.** The only prune-adjacent row is **R-210**, which is `WAITING-ON-OPERATOR`, says in terms *"Nothing was deleted; this is a list, not an action"*, and concerns **local Docker images on DooPlex**, not the Gitea registry | | `drives.enrol` | **RESTORED**`audits/SPIKE-raw-drive-enroll-2026-06-15.md`: a live drive walked scan → format → mount → PVE storage → one-click enrol, with the resulting `storage.cfg` entry and mount unit recorded |
| "newest ten per package" | **Not visible in the current state.** `felhom-agent` generic holds 10, but `felhom-controller` and `felhom-hub` container packages hold **19 each** | | `backup.tier1` | **RESTORED**`audits/CAMPAIGN-8-backup-restore-2026-07-27.md`: adversarial, destructive, unattended, both boxes + ep0; A2 proven end to end |
| `fail.lost-recovery-code` | **RESTORED**`audits/REHEARSAL-byo-reinstall-2026-08-09.md`, and proven the hard way on 2026-08-10 |
| `backup.restore-proof` | **grey — THE DECAY RULE FIRED** (below) |
| `install.installer-by-tag` | grey — hits are ISO spikes; nothing walks a tag rollback |
| `use.lifecycle` | grey — two passing mentions, no walk |
| `drives.migrate` | grey — one or two mentions only |
| `backup.whole-machine` | grey — diagnostics and phase findings, no walk of the claim |
| `fault.selfheal` | grey — **the best hit argues the other way**: `INCIDENT-guest-dhclient-killed-2026-07-20.md` documents self-heal's *absence* through a 1 h 15 m outage |
| `fault.operator-email` | grey — source-verified, delivery never observed |
| `fail.drive-filling` | grey — weak hits only |
| `fail.hub-down` | grey — Campaign 11's hub-unreachable work produced a *finding* (R-224), not a pass |
Everything else I could reach is silent, and silence here is not evidence of anything: ## 5. Did the "code moved under the proof" rule fire? **Yes — for the first time.**
`package_cleanup_rule` is **empty**; `package_version` has **no soft-delete column**, so a deletion
leaves no row; Gitea's `action` feed carries **no package operation at all** across 2026-08-08 →
2026-08-09 and nothing whatever on the evening of 08-08; the Gitea pod has **53 days uptime, 0
restarts**.
**And I withdrew one of my own claims.** R-287 said *"no DELETE on the packages API appears in 48 h of `backup.restore-proof` **has** a receipt: `architecture/_recovery-inventory-2026-07-28.md` carries live
Gitea router logs"*. Re-checked: `kubectl logs --since=72h` returns nothing older than **2026-08-09 journal lines for scheduled restore-tests on both boxes and both tiers. It is superseded anyway —
16:35** and contains **zero** `api/packages` lines even for requests I made myself. **The log never demo-hp logged `restore_test_failed` on 2026-08-05, and the box has since been wiped and reinstalled.
covered the window, so its silence was never evidence.** Withdrawn in the register. The claim is about a **continuing** scheduled behaviour, so a 2026-07-28 observation cannot carry it.
**What remains established, unchanged:** `v0.120.0` downloadable at 2026-08-08 14:29 UTC (run 267 **So the rule can fire, and now has.** Two nights ago it fired zero times out of twelve because every
printed `ok v0.120.0`), 404 by 2026-08-09 09:30 UTC (run 284), and `0.128.0` published at 14:47 UTC — downgrade came from *missing* evidence, not decayed evidence — there was nothing for it to bite on.
eighteen minutes after the good run. **It may be unestablishable from this side: Gitea keeps no **A consequence to note: the capability map still reads `PROVEN-LIVE (2026-08-03)` for that row, so the
package-deletion trail.** map is now the thing out of step**, and it is the source. Recorded rather than silently rewritten.
Per §8 I halted that part's attribution and carried on; the rest is independent. ## 6. Part 4 — the promise, and why the wording is what it is
## 2. The number, and the two places that read it **Surfaces:**
`felhom-agent/scripts/retention-policy.json``generic_versions_kept: 10`. | file:line | |
|---|---|
| `controller/internal/web/templates/backups_remote.html:101` | **the false promise** |
| `controller/internal/web/templates/backups_remote.html:98` | the orphan explanation — accurate, keep |
| `controller/internal/web/templates/layout.html:143` | the 14-day abandon countdown — accurate, keep |
| `controller/internal/settings/settings.go:337-338` | `OrphanedRenamedTo` schema comment — accurate, keep |
Read by **`scripts/check-published-versions.py`** (bounds its assertion) and referenced by the prune *(The capability map was searched and makes no such claim — nothing to correct there.)*
procedure. The file states, in its own header, that ten is an **observed state and not a located
ruling**, and that the principled bound is the hub's vouched `min_agent` floor — nothing can install
below it — which needs network the gate does not have.
**What CI no longer covers:** *a released version older than the retention window is no longer **The three cases:** set aside **before** the fix → not recoverable by construction (the restic
asserted downloadable.* Its **git tag and config tree are still asserted**; only the binary's presence password lives only in the identity bundle, `escrow/identity.go:39` read by `escrow/recover.go:91`, and
is dropped. The check **prints the dropped versions every run**: those rows are NULL); set aside **after** → recoverable in principle, never demonstrated; **today's
population is entirely the first case**.
``` **The deciding fact: neither the box nor the customer can tell which case they are in.** The hub holds
11 released version(s); retention policy keeps the newest 10 the discriminator (`host_escrow_superseded.identity_blob`); the box caches only
NOT ASSERTED (older than the retention window …): 0.120.0 `HubEscrowIdentityPresent`, which is about the *current* escrow; and **no field on the report or ACK
^ these versions still have git TAGS … what is no longer asserted is the BINARY's presence. wire carries superseded-blob retention**. The box renders the card. **A conditional promise the
``` renderer cannot evaluate is the same defect as an unconditional false one** — so the specified copy
stops promising, says plainly what happens, explains why it cannot promise, and names a route
(write to us).
Three controls: green at 10 naming what it dropped (**exit 0**); widened to 11 the evicted version **Specification:** `documentation/design/SPEC-orphan-card-copy-2026-08-10.md` — copy, surfaces, and the
re-enters and convicts (**exit 1**, `FAIL v0.120.0`); the policy file removed → **exit 2 INCONCLUSIVE**, render tests that should pin it, including a regression guard that the string `visszaállítható lehet`
naming the path it tried — never silently unbounded. never returns. **Not implemented, on purpose:** it lands in the controller, and a controller release is
undelivered until a golden carries it (R-242). It should ship with the next controller change so one
bake and one approval cover both, and the spec says so.
## 3. `felhom-agent` green ## 7. Register
``` **Ceiling R-292 → R-294.** Opened **R-293** (the census) and **R-294** (the promise + spec). **R-287
reuse-refs OK · instructions OK · published OK · release-complete OK · all agent gates OK turned to ESTABLISHED** with my unpaginated-count error withdrawn. **R-291's blocking condition
``` released.** **R-288 gained a second measured cost**, and it is a different failure mode from the first:
prose rows make claims ambiguous; rows-about-other-things make facts unfindable.
**At `main`: green**, pushed as `53d047a`, and the pre-push hook ran the same entry point. ## 8. Observations — noticed, not acted on
**At a tag: not re-proved tonight, and I will not claim it.** The available evidence is that the gate
is ref-independent — it enumerates from the Gitea tags API, and runs **190** (`v0.126.0`) and **216**
(`v0.127.0`) were tag pushes that passed. Minting a tag purely to prove it would have published a
release, which this session forbids.
## 4. Red-proofs - **The capability map is now out of step in two directions** — behind the dataset for three claims it
already called PROVEN-LIVE, and ahead of it for `backup.restore-proof`. Both point at R-288.
| # | mutation | asserted applied | outcome | - **`fault.selfheal`'s only real document argues against it.** Worth someone deciding whether the
|---|---|---|---| capability is real and unwalked, or overstated.
| 1 | **tag check removed** | `grep -c` → 1 | **scenario A FAILED, reporting `flash = "artifacts_set"`** — Friday's exact defect returned: the manifest saved with no tag | - **Retention is proven; recovery from a superseded row is not.** The census proves blobs are now kept;
| 2 | package check removed | `grep -c` → 1 | scenario B FAILED, `flash = "artifacts_set"` | nobody has ever unwrapped one. That is the next thing worth a drill, and it needs no customer.
| 3 | inconclusive branch mapped onto success | `grep -c` → 2 | scenario E FAILED | - The page now shows the three restored claims as **"moved twice"** rather than once, so an unsettled
status reads as unsettled.
All reverted; `grep -c MUTATION`**0**; all seven tests green again.
**Red-proof 3, stated precisely rather than flatteringly:** with the inconclusive branch mapped onto
success the save did **not** complete — it fell through to `artifact_sha_invalid`. So the mutation
proves the guard is load-bearing for *the message the operator sees*, not for the save itself. That is
the honest reading, and it is exactly the defect scenario E exists to prevent: the operator being told
"missing or invalid" when the truth is "I could not reach the registry".
## 5. The five scenarios, as the operator sees them
| | outcome | the message |
|---|---|---|
| **A** tag missing | REFUSED | *"Refused: that version has no usable git tag. The installer fetches an agent's config files from `raw/tag/v<version>/configs/`, so a version published without its tag makes every fresh install and reinstall fail at step 5 of 8 — as root, on a virgin machine. … Fix it by pushing the tag: `git tag -a v<version> <released-commit> && git push origin v<version>`"* |
| **B** package pruned | REFUSED | *"Refused: that version's artifact is not downloadable. The version is tagged but its package is not in the registry, so a box would 404 fetching the binary itself. … Publish it — `bash scripts/release-agent.sh <version>`"* |
| **C** no checksum | REFUSED | the pre-existing *"Couldn't set the checksum …"* — see the note below |
| **D** both good | **SAVES**, `artifacts_set`, byte-identical behaviour | — |
| **E** registry unreachable | REFUSED | *"Refused: could not verify — this does not mean anything is missing. … It refuses rather than saving with a warning, because a warning beside a success reads as a success. … There is deliberately no override: the registry is on your own server, so if it is unreachable the vouch can wait."* |
**Scenario C changed shape because my first draft modelled nothing real,** and that is worth recording.
With a Gitea client configured, `resolveArtifactSHA` fetches the sha **authoritatively and ignores what
was submitted** — so "submit an empty sha" cannot produce an empty stored sha. The genuine shape is
Gitea answering with no `sha256`, and the **existing** refusal already owns it. The test now pins the
guarantee (*the manifest is unchanged*) rather than a mechanism I had invented.
**No override was built, and none is wanted.** §8's halt condition did not trigger.
## 6. Part 4 — NOT DONE
The twelve downgraded claims were not re-examined and no receipts were searched for. **The honest-grey
count is therefore still the twelve from last night, unverified in either direction**, and
`where-felhom-stands.*` is untouched. This is the session's shortfall: Parts 13 took the budget, and
splitting Part 4 in half would have produced exactly the kind of half-checked green the whole exercise
exists to prevent.
**Whether the "code moved under the proof" rule can fire at all** is therefore still open from last
night, where it fired **zero** times out of twelve — every downgrade came from missing evidence, not
from decayed evidence. My reading remains that it *can* fire but will stay rare until rows cite
evidence at all, which is R-290.
## 7. Hub deployed
**v0.102.0**, live and verified: deploy image `gitea.dooplex.hu/admin/felhom-hub:0.102.0`, rollout
complete, and the page footer reads `0.102.0`. Manifest commit **`36bcd12`**; code commit `b55fc17`.
The image was verified **served by the registry before** the manifest was bumped, not after. ArgoCD's
`felhom` app has `automated.enabled: false`, so the sync was triggered explicitly — **no
`kubectl set image` at any point.**
## 8. Part 5 — dropped, as marked
Not started. It was explicitly droppable and it is the only part that touches customer-facing wording,
which is where a rushed edit does most harm.
## 9. Register
**Ceiling R-290 → R-292.** Opened **R-291** (the narrowing, with its reason, so it can be widened
deliberately) and **R-292** (`artifact_sha_invalid` conflates three facts — found by scenario E
failing, worked around by ordering, message untouched). Closed: **R-273's owed-guards tail**, both
guards built. Corrected: **R-287**, twice.
## 10. Observations — noticed, not acted on
- **R-292 is the interesting one.** A test I wrote to check a new guard failed for a reason that had
nothing to do with the guard, and that reason was a real pre-existing defect. The five scenarios
earned their keep before the feature shipped.
- **The golden gets no tag probe.** It is fetched by version and has no config tree, so a tag probe
would assert something the installer never does. Deliberate, and stated in the code.
- **The gate is skipped entirely when no Gitea client is configured**, or a hub without registry
credentials could never vouch anything. Pinned by its own test, and worth knowing: the guard is only
as present as the client is.
- `check-release-complete` runs in `--fast`, so a missing tag is now caught by the **pre-push hook**,
earlier than CI.
+56 -87
View File
@@ -27,6 +27,52 @@ So I took your stated fallback: the old store was **moved aside, not deleted** (
**One thing that needs your judgement, not mine.** The card that offered this told the customer their set-aside backups *may be restorable later with their recovery code*. For these ones that is simply untrue, and it is said to precisely the people who have just lost their history. *(R-202 — now evidenced.)* **One thing that needs your judgement, not mine.** The card that offered this told the customer their set-aside backups *may be restorable later with their recovery code*. For these ones that is simply untrue, and it is said to precisely the people who have just lost their history. *(R-202 — now evidenced.)*
## Can anyone else lose their history the way demo-felhom did? No.
**One read of the hub's own records, no machine touched.** The hub holds backup keys for exactly
**three** machines. Both demo boxes lost their old key in the same four-hour window on 4 August,
before the retention fix was in force — that is the whole population of the problem, and it is
entirely ours. **The tester's machine has no record at all**, so it cannot be affected; and anything
enrolled from now on is covered, because the fix has been in force since 4 August.
I ran a control before trusting the query: it had to say *material present* for a machine known to
have it and *absent* for one known not to. It did both.
## Three green dots came back, nine stayed grey, and one rule finally fired
**The nine greys are the honest number.** For those, no document anywhere walks the claim, and saying
so is more useful than a dot nobody can defend.
- **Back to green**, each citing the document that walked it: the drive wizard (a live drive taken
through scan → format → mount → enrol), the on-box app backups (an overnight destructive campaign
across both machines), and the lost-recovery-code case — which we then proved the hard way this
morning.
- **One claim stayed grey for a new reason, and it is the interesting one.** The unattended
restore-proof *does* have a receipt from 28 July — but demo-hp's restore-test failed on 5 August and
the machine has since been wiped and rebuilt. It is a claim about something that keeps happening, so
an old observation cannot carry it. **This is the first time that rule has fired**; two nights ago it
fired zero times out of twelve.
## The prune mystery is solved, and the answer was written down all along
Who deleted the old versions: **you did, on 4 August evening, on your own rule** — 33 deletions, keep
set asserted first, every one a clean 204. **It was recorded inside the row about the Configuration
page being slow**, because pruning artifacts is what made that page fast. Two sessions failed to find
it. **You are no longer blocked** on establishing something that was already on file.
I also got a number wrong yesterday and it is corrected: I said the container packages held nineteen
versions and used that to argue against the prune. Counted properly — with pages — they hold 270 and
169, and the two that *were* pruned sit at exactly ten each.
## The sentence we should stop saying
When a machine's off-site history is set aside, the card tells the customer it *may be restorable
later with their recovery code*. **The machine showing that card cannot know whether it is true**
the fact lives on the hub and is not sent to the box. For anything set aside before 4 August it is
simply false. **The replacement wording is written and waiting**
(`documentation/design/SPEC-orphan-card-copy-2026-08-10.md`); it ships with the next controller
release so one image bake and one approval cover it, rather than costing you two of each.
## What works ## What works
A blank machine boots the Felhom disc, installs itself unattended, and is claimed by the customer, who A blank machine boots the Felhom disc, installs itself unattended, and is claimed by the customer, who
@@ -40,74 +86,18 @@ destroyed on purpose and its files came back byte for byte identical — four ti
their recovery code got everything back with **no command line inside the machine at any point**, their recovery code got everything back with **no command line inside the machine at any point**,
in 72 seconds. The two rough edges that walk found are also gone. *(R-201, R-252, R-253 — closed.)* in 72 seconds. The two rough edges that walk found are also gone. *(R-201, R-252, R-253 — closed.)*
## The guard is in. An approval that cannot be installed is now refused. ## Shipped 2026-08-09 — the guards, and the rehearsal that earned them
**Yesterday morning every install in existence failed for hours, and nothing would have stopped it **An approval that cannot be installed is now refused** at the moment you press Save: the hub checks
happening again. Now something does.** When you press Save on the Day-0 artifacts, the hub checks — the version's git label and that its file downloads, and refuses with a message naming the fix. A
before it writes — that each version you are vouching actually has its git label **and** that its file second machine catches it a step earlier in the agent repo. Both were owed after every install in
can actually be downloaded. If either is missing it refuses and tells you which, for which version, existence failed for hours on 2026-08-09. **Hub v0.102.0 is live.**
and the one command that fixes it. **Hub v0.102.0 is live.**
Three details worth your knowing: **The reinstall rehearsal:** a demo machine was wiped and put back. All four test files returned
- **It checks the exact file the installer fetches first** — the one whose absence broke Friday — not **byte for byte**, accented Hungarian filenames included, checked as raw bytes. But it only finished
some other file that happens to exist. A test pins that, because probing the wrong file is precisely because a terminal was available twice — the install died on a missing version label *(R-273, now
how the failure stayed invisible. guarded)* and **a reinstalled machine still cannot re-attach its own data drive** *(R-280 — the one to
- **"Could not check" also refuses**, with a different message. Saving with a warning would read as a fix before the tester's visit)*. Detail: `documentation/audits/REHEARSAL-byo-reinstall-2026-08-09.md`.
success, and we have the scars. **There is no override**: the registry is on your own server, so if
it is unreachable the approval can wait.
- **A second machine now catches it one step earlier** — the agent repo refuses to consider a release
complete unless its label and its package both exist.
## The red repository is green, and the two rules now share one number
The tidy-up keeps the newest ten versions; the check demanded every version ever labelled still be
downloadable. Both are sensible and together impossible, so the red would have returned on your next
publish. They now read the same number from one file. **What the check no longer covers, plainly: a
version older than the ten is no longer asserted downloadable** — its label and its config files still
are — and it prints which ones it dropped on every run so this cannot go quiet.
**One thing I could not establish, and I am not guessing.** Who actually deleted the old versions is
**still unknown**. Gitea keeps no deletion trail: no cleanup rule is configured, the version table has
no deleted-marker, the activity feed shows no package operation, and the server log no longer reaches
back that far. I also **withdrew my own claim from yesterday** that the logs showed no deletion — the
logs did not cover the window, so they never said anything.
## The rehearsal finished. The data came back byte for byte; the journey did not.
**We wiped a working demo machine and put it back. All four test files returned identical — including
the two with Hungarian accents, checked as raw bytes, not as text on screen.** The unlock took 21
seconds and the restore 13. **But it only finished because I could open a terminal twice.** A
household would have stopped, twice, and the second time the screen would have told them it was easy.
**The two walls, both fixed-or-fixable, neither about the data:**
- **The install died four steps in** — the agent version you approved had been published as a download
but never given its version label, and the installer looks it up by that label. **Now unblocked**
I pushed the label after checking the published file matched what you vouched. *(R-273 — closed. The
two guards that would stop it recurring are still owed.)*
- **A reinstalled machine cannot re-attach its own data drive.** Every route is a dead end, and the
restore page cheerfully says „**Ez két kattintás**" while pointing at an empty list. The drive is
fine and the machine can see it — it just is not offered, because the same drive is also the backup
target. I got past it by typing an internal path no customer could know. **This is the one to fix
before the tester's visit.** *(R-280)*
**Also broken, found on the way:**
- **Taking Felhom off a machine leaves the one thing that stops it going back on.** We install a small
network service at setup; removing Felhom restarts it without its settings, it seizes the port the
next install needs, and the next install then refuses — appearing to blame the owner's network.
*(R-272)*
- **A machine we removed keeps its private line to us open.** *(R-276)*
- **The hub said nothing at all** while a machine was wiped, rebuilt, re-claimed and had its sealed
backups opened. No false alarm — but also no word, and the alarm that exists for "someone is opening
this customer's backups" stayed silent through a real one. *(R-281)*
- **A rebuilt machine may still come back on software from last week** — narrower than I first wrote:
the resumed install fetched the right version, but a fresh one takes whatever copy is newest on the
disk without checking it against what you approved. *(R-274)*
- **demo-felhom has not had an off-site backup in six days** and is waiting for a recovery code nobody
has entered. demo-hp, rebuilt the same day, recovered by itself. *(R-278)*
- **One code, three different names**, and the email points at a page the machine is not showing —
this cost us a wasted code today. *(R-282, R-283)*
## What's broken ## What's broken
@@ -121,27 +111,6 @@ household would have stopped, twice, and the second time the screen would have t
accumulates. *(R-244)* accumulates. *(R-244)*
- **Putting restored files back where they belong is still manual.** *(R-213)* - **Putting restored files back where they belong is still manual.** *(R-213)*
## The rest of what the rehearsal found
Sixteen findings in one afternoon, **none of them visible from reading the code** — three sessions of
review had not seen any.
- Removing Felhom leaves five files holding old keys *(R-275)*, and rotating a leaked key does not
revoke the old one until the service restarts *(R-269)* — the written recipe for it is a step short
*(R-270)*, and the alarm it raises can never be closed because the fix it recommends is what
silences the all-clear *(R-271)*.
- **It caught me being wrong twice, and that matters more than the count.** I told you the fleet's
off-site backups were down; demo-hp had eighteen snapshots and I had read three misleading screens
instead of asking the machine *(R-277)*. And I raised a leftover permissions file as a security
hole, then tested it and refuted myself — it is inert.
**What worked, and should not be lost in the count:** the machine came up on its own at the approved
version; the setup page appeared unprompted, in Hungarian, naming the customer; **the recovery screen
appeared without being looked for** and said plainly that unlocking changes nothing; the restore told
the truth about putting files in a checking folder rather than back in place; and no false alarm fired.
Full account: `documentation/audits/REHEARSAL-byo-reinstall-2026-08-09.md`.
## Three rulings, written down so they stop living in a conversation ## Three rulings, written down so they stop living in a conversation
- **The managed-update floor.** It is deliberately parked, and the trigger to raise it is **the first - **The managed-update floor.** It is deliberately parked, and the trigger to raise it is **the first
File diff suppressed because one or more lines are too long
@@ -240,17 +240,20 @@ claims:
band: journey band: journey
stage: 4 stage: 4
title: "A new drive is found, offered, formatted, mounted and enrolled — including on awkward older boot layouts" title: "A new drive is found, offered, formatted, mounted and enrolled — including on awkward older boot layouts"
status: built status: walked
note: "Applies to a NEW drive. Re-attaching an existing one after a reinstall is R-280 and fails." note: "Applies to a NEW drive. Re-attaching an existing one after a reinstall is R-280 and fails."
sources: sources:
- capability-map: "Drive wizard: scan/format/mount/enroll, incl. legacy-boot LVM-root hosts" - capability-map: "Drive wizard: scan/format/mount/enroll, incl. legacy-boot LVM-root hosts"
- register: "R-220" - register: "R-220"
- evidence: "audits/SPIKE-raw-drive-enroll-2026-06-15.md"
changed: changed:
from: walked from: built
reason: "the 2026-08-09 walk exercised RE-attach (which failed, R-280); first-enrolment of a NEW drive has no walk on file" also_moved: 2026-08-09 walked -> built
reason_superseded: "the 2026-08-09 walk exercised RE-attach (which failed, R-280); first-enrolment of a NEW drive has no walk on file"
reason: "receipt found 2026-08-10: a live throwaway drive on felhom-pve walked scan -> format -> mount -> PVE dir storage -> one-click Regisztralas, with the resulting storage.cfg entry and systemd mount unit recorded. The map already read PROVEN-LIVE; the dataset was behind it. Scope unchanged: this is a NEW drive. Re-attaching an existing one after a reinstall is R-280 and still fails."
verified: verified:
date: 2026-08-09 date: 2026-08-09
verdict: downgraded verdict: upgraded
depth: source-read depth: source-read
- id: drives.migrate - id: drives.migrate
band: journey band: journey
@@ -297,16 +300,19 @@ claims:
band: journey band: journey
stage: 5 stage: 5
title: "App data on the machine, nightly database dumps, a copy on a second drive" title: "App data on the machine, nightly database dumps, a copy on a second drive"
status: built status: walked
sources: sources:
- capability-map: "Tier-2 secondary-drive copy: class-driven legs" - capability-map: "Tier-2 secondary-drive copy: class-driven legs"
- evidence: "audits/CAMPAIGN-8-backup-restore-2026-07-27.md"
changed: changed:
from: walked from: built
reason: "no walk document cited" also_moved: 2026-08-09 walked -> built
reason_superseded: "no walk document cited"
reason: "receipt found 2026-08-10: an adversarial, destructive, unattended overnight campaign across both boxes and ep0, with A2 (one quiesce, two tiers) proven end to end. Map already read PROVEN-LIVE."
verified: verified:
date: 2026-08-09 date: 2026-08-09
verdict: downgraded verdict: upgraded
depth: register+map depth: source-read
- id: backup.whole-machine - id: backup.whole-machine
band: journey band: journey
stage: 5 stage: 5
@@ -348,6 +354,7 @@ claims:
changed: changed:
from: walked from: walked
reason: "no walk document cited, and the last recorded restore-test on demo-hp FAILED (2026-08-05)" reason: "no walk document cited, and the last recorded restore-test on demo-hp FAILED (2026-08-05)"
decay: "PROOF-DECAY RULE FIRED (first time it has). A receipt EXISTS - architecture/_recovery-inventory-2026-07-28.md carries live journal lines for scheduled restore-tests on both boxes and both tiers - but it is superseded by later observation: demo-hp logged restore_test_failed on 2026-08-05, and the box has since been wiped and reinstalled (2026-08-09). The claim is about a CONTINUING scheduled behaviour, so a 2026-07-28 observation cannot carry it. Stays grey until a scheduled restore-test is seen passing on the rebuilt box. THE CAPABILITY MAP STILL READS PROVEN-LIVE (2026-08-03) AND IS NOW THE THING OUT OF STEP."
verified: verified:
date: 2026-08-09 date: 2026-08-09
verdict: downgraded verdict: downgraded
@@ -621,17 +628,20 @@ claims:
- id: fail.lost-recovery-code - id: fail.lost-recovery-code
band: failures band: failures
title: "The customer loses their recovery code — by design, the data is unrecoverable" title: "The customer loses their recovery code — by design, the data is unrecoverable"
status: built status: walked
note: "The recovery screen states it in Hungarian." note: "The recovery screen states it in Hungarian."
sources: sources:
- capability-map: "Escrow ceremony: customer-facing wizard, one-shot R claim, operator zero-knowledge" - capability-map: "Escrow ceremony: customer-facing wizard, one-shot R claim, operator zero-knowledge"
- register: "R-198" - register: "R-198"
- evidence: "audits/REHEARSAL-byo-reinstall-2026-08-09.md"
changed: changed:
from: walked from: built
reason: "a by-design refusal; no walk document cited" also_moved: 2026-08-09 walked -> built
reason_superseded: "a by-design refusal; no walk document cited"
reason: "receipt found 2026-08-10, and then proven the hard way: the recovery screen states in Hungarian that nobody can replace a lost code, and on 2026-08-10 demo-felhom's pre-fix key was confirmed unrecoverable by construction (identity_blob NULL; the restic password lives only in that bundle). The by-design refusal is real."
verified: verified:
date: 2026-08-09 date: 2026-08-09
verdict: downgraded verdict: upgraded
depth: source-read depth: source-read
- id: fail.moves-house - id: fail.moves-house
band: failures band: failures
File diff suppressed because one or more lines are too long
@@ -0,0 +1,90 @@
# SPEC — the orphan card must stop promising what it cannot know (2026-08-10)
**Specification only. Nothing here is implemented, deliberately** — see §5.
**Why this exists.** When a machine's off-site history is set aside, the card tells the customer — in
Hungarian, at the moment they have just lost that history — that the old copies *may be restorable
later with their recovery code*. **For anything set aside before 2026-08-04 that is false and
unfixable**, and on 2026-08-10 it was said to a real machine in exactly that state (`demo-felhom`).
Register: **R-202**.
---
## 1. The surfaces, at `file:line`
All in `felhom-controller`:
| # | file:line | what it says today |
|---|---|---|
| 1 | `controller/internal/web/templates/backups_remote.html:98` | *"A távoli tárhelyen lévő mentések egy korábbi, már nem elérhető kulccsal készültek…"* — the orphan explanation. **Accurate; keep.** |
| 2 | `controller/internal/web/templates/backups_remote.html:101` | *"A régi előzmény **félretéve marad** (nem törlődik), és a hozzá tartozó helyreállítási kóddal később visszaállítható lehet."* — **the false promise.** |
| 3 | `controller/internal/web/templates/layout.html:143` | the 14-day abandon countdown — *"a korábbi távoli mentéseidet {{.RecoveryAbandonDays}} nap múlva véglegesen töröljük"*. **Accurate; keep**, but see §3. |
| 4 | `controller/internal/settings/settings.go:337-338` | schema comment on `OrphanedRenamedTo`: *"so the card/log can name where the old history was set aside."* **Accurate; keep.** |
The `00-capability-map.md` restorability wording was searched for and **not found** — the map does not
make this claim, so there is nothing to correct there. (Stated because the task expected one.)
## 2. What is actually true, in three cases
| case | truth |
|---|---|
| material set aside **before** hub v0.93.0 (in force 2026-08-04 ~11:11Z) | **Not recoverable, by construction.** The restic password lives only in the identity bundle (`escrow/identity.go:39`, read by `escrow/recover.go:91`); for these rows `identity_blob` is NULL. No code and no recovery code can produce it. |
| material set aside **after** the fix | **Recoverable in principle** — the blob is retained and the customer's code unwraps it. Never yet demonstrated end to end on a superseded row. |
| **today's population** | Both cases that exist are the first one. Census 2026-08-10: two hosts hold a supersession, both `identity_blob` ABSENT, both superseded ~4 h before the fix. |
## 3. THE DECIDING FACT: neither the box nor the customer can tell which case they are in
- **The hub can tell** — it holds `host_escrow_superseded.identity_blob` (`store/store.go:393`).
- **The box cannot.** It caches only `HubEscrowIdentityPresent`
(`controller/internal/settings/settings.go:71`), which is about the **current** escrow, not a
superseded one. **No field on the report or ACK wire carries superseded-blob retention.**
- **The box is what renders the card.**
**So the card is making a conditional promise its renderer cannot evaluate. That is the same defect
as an unconditional false promise, in a longer sentence.** The wording below therefore stops
promising and starts stating.
## 4. The replacement copy
**Surface 2 — `backups_remote.html:101`. Replace:**
> A régi előzmény **félretéve marad** (nem törlődik), és a hozzá tartozó helyreállítási kóddal később
> visszaállítható lehet.
**with:**
> A régi előzmény **félretéve marad a tárhelyen — nem töröljük**. Új, üres tárolót hozunk létre, és a
> következő mentés oda készül.
>
> **A félretett mentések megnyithatóságát itt nem tudjuk megígérni.** Ez attól függ, megvan-e még a
> hozzájuk tartozó kulcs, és ezt ez a gép nem tudja megállapítani. Ha szeretnéd, hogy utánanézzünk,
> **írj nekünk** — a félretett másolat addig is a helyén marad.
Three properties, each deliberate: it **states** what happens (set aside, not deleted, fresh store);
it **declines** the claim it cannot evaluate, and says *why* rather than going vague; and it **names a
route the customer can take** — writing in — instead of leaving them nowhere.
*(Surface 3's countdown stays, but the two must be read together: the card must not simultaneously say
"we cannot promise these are openable" and "we will delete them in N days" without the second making
clear it is the set-aside copy being deleted. Recommend re-reading the pair once implemented.)*
## 5. The test that should pin it
`controller/internal/web/` render test, one per branch of the orphan card:
1. **Orphan card rendered → the string `visszaállítható lehet` does not appear anywhere in the
output.** That is the regression guard: it fails if the promise returns in any form.
2. **Orphan card rendered → the "write to us" route is present.** A refusal that names no route is a
defect in this project; pin the route, not only the absence.
3. A render test **per branch of the gate** that shows the card at all, per the seam-wiring rule —
a handler test that POSTs proves nothing about reachability.
## 6. Why it is not implemented tonight, and what it should ship with
The change lands in `felhom-controller`, and **a controller release is not delivered until a golden
carries it** — which costs one image bake and one operator approval (R-242 is the row about exactly
that gap). **The next session changes the controller anyway.** One bake and one approval should cover
this copy change together with that work, rather than spending two of each.
**Do not let this decouple:** if the next controller release ships without this, the promise is still
being made to the next customer who loses a history.
+13 -2
View File
@@ -93,9 +93,16 @@ def claim_html(c):
body = ['<span style="font-size:12px;line-height:1.45;color:#b8c4d6;text-wrap:pretty">', body = ['<span style="font-size:12px;line-height:1.45;color:#b8c4d6;text-wrap:pretty">',
esc(c.get("title", ""))] esc(c.get("title", ""))]
if c.get("changed"): if c.get("changed"):
ch = c["changed"]
if ch.get("also_moved"):
# Moved twice. Say so — a second move rendered as a first one hides the fact that a
# status is unsettled, which is exactly what a reader needs to see.
label = "moved twice: %s, then 2026-08-10 back to %s" % (
esc(ch["also_moved"]), esc(c.get("status", "?")))
else:
label = "changed 2026-08-09, was %s" % esc(ch.get("from", "?"))
body.append('<span style="display:inline-block;margin-left:6px;padding:1px 5px;border-radius:3px;' body.append('<span style="display:inline-block;margin-left:6px;padding:1px 5px;border-radius:3px;'
'background:#3a2a12;color:#fbbf24;font-size:10px;white-space:nowrap">changed 2026-08-09, was %s</span>' 'background:#3a2a12;color:#fbbf24;font-size:10px;white-space:nowrap">%s</span>' % label)
% esc(c["changed"].get("from", "?")))
if c.get("v_verdict") in ("needs-hardware", "contested"): if c.get("v_verdict") in ("needs-hardware", "contested"):
col = "#a78bfa" if c["v_verdict"] == "contested" else "#7dd3fc" col = "#a78bfa" if c["v_verdict"] == "contested" else "#7dd3fc"
body.append('<span style="display:inline-block;margin-left:6px;padding:1px 5px;border-radius:3px;' body.append('<span style="display:inline-block;margin-left:6px;padding:1px 5px;border-radius:3px;'
@@ -104,6 +111,10 @@ def claim_html(c):
if c.get("note"): if c.get("note"):
body.append('<span style="display:block;color:#7c8aa3;font-size:11px;margin-top:3px">%s</span>' body.append('<span style="display:block;color:#7c8aa3;font-size:11px;margin-top:3px">%s</span>'
% esc(c["note"])) % esc(c["note"]))
decay = c.get("decay") or (c.get("changed") or {}).get("decay")
if decay:
body.append('<span style="display:block;color:#c08a3e;font-size:11px;margin-top:2px">'
'proof decay: %s</span>' % esc(decay))
if c.get("changed", {}) and c["changed"].get("reason"): if c.get("changed", {}) and c["changed"].get("reason"):
body.append('<span style="display:block;color:#9a7b3a;font-size:11px;margin-top:2px">why it moved: %s</span>' body.append('<span style="display:block;color:#9a7b3a;font-size:11px;margin-top:2px">why it moved: %s</span>'
% esc(c["changed"]["reason"])) % esc(c["changed"]["reason"]))