Verify the standing picture against source: 12 downgrades, and the decay ran both ways
gates / gates (push) Successful in 21s

55 claims verified. Twelve moved, all downwards: walked 32 -> 20, built 5 -> 17.
Register ceiling R-284 -> R-290.

THE RULE DID NOT FIRE THE WAY IT WAS EXPECTED TO. Not one downgrade came from
code moving under an old proof. All twelve came from step 1 of the same rule --
the cited evidence does not exist. Measured: of the 28 capability-map rows
behind the page's claims, 8 carry a tests/ or audits/ path and 20 carry prose
only. The green dots were drawn from rows that cite an argument, not a walk
(R-290). The map, not the dataset, is what needs fixing -- it still says
PROVEN-LIVE for all twelve.

And once it ran backwards: fault.operator-email looked contradicted by R-182,
but live source shows the backup_run_failures digest allowlisted, operator-only
and templated, with recovery_unit_capture_failed now record-only. The claim is
right and the REGISTER ROW is stale (R-289). The session went looking for stale
proofs and found a stale defect.

R-281 WITHDRAWN -- wrong in both directions, settled by the operator's mailbox.
The tripwire DID fire (escrow_blob_served 10:19:41Z = 12:19 CEST) and false
error-severity alarms fired too, for deliberate attended work (R-285). The
measurement's cause is ESTABLISHED: the P7 query copied hub.db without hub.db-wal,
and the signature is exact -- it reported "2 events all day, newest 00:30:07",
and the rows at or before 00:30:07 number exactly 2. Timezone and wrong-key were
tested and refuted. The control had been drawn from the same stale snapshot as
the measurement, which is why it agreed (R-286).

Part 4: NO WORKFLOW CHANGED, deliberately. The gate is not ref-sensitive -- it
enumerates from the Gitea tags API, and both previous tag pushes passed. The red
is TRUE: run 267 saw v0.120.0 downloadable, run 284 on the same commit saw 404.
Who deleted the package is NOT established and is not guessed (R-287).

The page is now generated from where-felhom-stands.yaml by scripts/render_stands.py:
static, zero script tags, every moved status carrying a visible "changed, was X"
chip. The React bundle -- whose content was gzip+base64 inside a JS module map --
is kept as a dated snapshot. scripts/check_stands.py gates the data and convicted
51 problems in my own first draft before the staged positive control ever ran.
This commit is contained in:
2026-08-09 18:40:49 +02:00
parent a199c492f4
commit 6088afcbed
8 changed files with 1274 additions and 148 deletions
+154 -147
View File
@@ -1,184 +1,191 @@
# REPORT — the seed that never ran twice, and three pictures that were not true (2026-08-08)
# REPORT — making the picture true (2026-08-09, unattended)
Four defects of one family, each with a source-verified mechanism, tests and red-proofs.
Agent **v0.128.0** · controller **v0.210.0** · `gates.yml` (workflow only). **The hub was not touched,
not bumped and not deployed.**
Read-only against all live infrastructure. **Both demo machines were powered off and in transit; no
box was probed, woken or waited on.** Claims that only a running box could settle are marked
`needs-hardware`, which is a verdict, not a gap.
## 1. Part 1's live sequence — RUN, OPERATOR-PRESENT, PASSED
---
On `demo-felhom` guest 9201, agent 0.128.0. Backup taken first and verified byte-identical
(`sha256 eff18437…`), pbsdr marker sha recorded and **unchanged at every step**
(`fd4c97b5…`).
## 1. The verdict table
**Step 3 — the preflight with the key removed (today's defect, live):**
**55 claims. Statuses moved on 12 of them — all downwards.** `walked 32 → 20`, `built 5 → 17`;
`partial 14` and `missing 4` unchanged. Full per-claim detail with sources is in
`documentation/architecture/where-felhom-stands.yaml`.
```
{"id": "pbs_storage_id", "ok": false, "detail": "escrow.pbs_storage_id not configured"}
```
**Step 4 — one 60 s tick later, same daemon:**
```
{"id": "pbs_storage_id", "ok": true, "detail": "felhom-pbs"}
```
**No restart:** agent MainPID **1993397 before and 1993397 after** — the same process. The wizard
would have refused at 18:04:16Z and would pass at 18:05:44Z, with nobody touching the box.
**Step 5 — nothing else moved:** 45 keys in the backup, 45 now, **added none / removed none /
changed none**. The one key is back with its original value.
**Positive control:** the preflight was exercised BEFORE the change and reported `ok: true` on the
same row, so a green afterwards is a measurement and not an artefact of the probe.
## 2. The writer of `agent.json` — ESTABLISHED
`step_agent_config`, **`felhom.eu/scripts/felhom-host-install.sh:2396`**; the Python render at
**`:2449`**; the `O_TRUNC` write at **`:2579`**. `PRESERVE_FROM` defaults empty (**`:256`**) and is set
only by an explicit `--preserve-from` (**`:1246`**). **The render never writes an `escrow` section at
all** — grep over the whole heredoc: zero hits. The pbsdr marker is host-side
(`<agent-state>/pbsdr/marker.json`) and survives. **R-221's attribution was correct.**
A rebuild is only the case that was *measured*; the same hole opens for a hand-edited or restored
config, which is the honest reason the fix is at the seam rather than in the installer.
## 3. Red-proofs — 6 of 6, each with the mutation asserted applied
| # | mutation | assertion it applied | outcome |
| claim | was | now | why |
|---|---|---|---|
| 1 | **Part 1 / Scenario A: remove the new seed call** | marker `MUTATED: the R-221 re-assert removed` present | **RED** — and **yes, it failed against today's tree**, with the intended message ✔ |
| 2 | Part 1: remove the early return as well | marker `MUTATED: early return deleted` present | **RED** — the zero-Proxmox-calls assertion is load-bearing, not decorative ✔ |
| 3 | Part 3 / F: revert to `status.LastDBDump.Success` | marker present | **RED** — app X's false green returns ✔ |
| 4 | Part 3 / G: map "no result" to `ok` | marker present | **RED** — green-on-presence returns ✔ |
| 5 | Part 2 / D: ignore `DiskKnown` in the template | marker present | **RED** — „0.0 GB / 0.0 GB (0%)" in the nominal colour returns ✔ |
| 6 | Part 2 / E: force the flag false | marker present | **RED** — a healthy box is shown losing its numbers ✔ |
| `install.installer-by-tag` | walked | **built** | gate 6 asserts the manifest names an installer tag; no walk of a rollback on file |
| `use.lifecycle` | walked | **built** | no walk document cited |
| `drives.enrol` | walked | **built** | the 08-09 walk exercised RE-attach (which failed, R-280); first enrolment of a NEW drive has no walk |
| `drives.migrate` | walked | **built** | no walk document cited |
| `backup.tier1` | walked | **built** | no walk document cited |
| `backup.whole-machine` | walked | **built** | no walk document cited |
| `backup.restore-proof` | walked | **built** | no walk cited, **and the last recorded restore-test on demo-hp FAILED** (2026-08-05) |
| `fault.selfheal` | walked | **built** | no walk document cited |
| `fault.operator-email` | walked | **built** | source-verified as correct, but no run observed delivering |
| `fail.drive-filling` | walked | **built** | no walk document cited |
| `fail.lost-recovery-code` | walked | **built** | by-design refusal; no walk document cited |
| `fail.hub-down` | walked | **built** | no walk document cited |
All six restored and re-verified green. **Answer to the question asked directly: the Part 1 test DID
fail against today's tree.**
**Upgraded: 1.** `install.byo` — the page said *"the first real one has not happened"*. A real
`--mode byo` install completed on demo-hp on 2026-08-09 (`Day-0 provision SUCCESS`, 3 m 49 s). Still
not a customer's own hardware, so not *walked*, but the sentence was false.
## 4. The Hungarian strings as shipped
**`needs-hardware`: 4** — `use.lan-fallback`, `backup.restore-proof`, `fail.disk-failing`,
`fail.internet-down`. Each needs an observation on a running box; each says which.
- `„A tárhely mérete most nem olvasható ki."` — the disk caveat line
- `„nem ismert"` — the short label in the value slot
- `„Erről a mentésről nincs eredményünk."` — the `title` on the no-verdict backup mark
**Confirmed: 38.** Seven of those were re-confirmed against live source or the live hub tonight rather
than against paperwork: the tripwire, the off-site repository, the claim path, the catalogue, the
tunnel, the reset code and the operator-email digest.
## 5. The §7.3 truth table as implemented
## 2. Every downgrade, with the coupling that broke
| this app's own most recent dump result | restore point | verdict |
|---|---|---|
| any of its databases failed | yes | `error` |
| all clean | yes | `ok` |
| none recorded | yes | **no icon**, time only, with the title above |
| any | no | no tier-1 row, unchanged |
The rule is *"a proof is about the code that existed when it ran"*. **It did not fire the way the task
expected.** Not one downgrade came from code moving under an old proof. **All twelve came from step 1
of the same rule — the cited evidence does not exist.**
**Recency was left alone**, deliberately: an age threshold means inventing a number, and the time is
already printed beside the icon. Recorded as an observation.
Measured: of the 28 capability-map rows behind the page's claims, **8 carry a `tests/` or `audits/`
path in their evidence column and 20 carry prose only.** The green dots were being drawn from rows
that cite an argument, not a walk. Filed as **R-290**.
## 6. The CI timeout, and what is still unknown
**And the decay ran the other way once.** `fault.operator-email` — *"one mail per run, every failing
app named"* — I first took to be contradicted by R-182 (open, *"tells the operator about ONE app and
silently swallows every other"*). Reading live source: the digest `backup_run_failures` is allowlisted
(`hub/internal/api/handler.go:1837`), operator-only (`notify/dispatcher.go:423`) and templated
(`notify/templates.go:48`); `recovery_unit_capture_failed` is record-only (`dispatcher.go:376`); a
cooldown drop now logs a `suppressed` row (`dispatcher.go:314-330`). **The claim is right and the
register row is stale** — filed as **R-289**. The session went looking for stale proofs and found a
stale defect.
**`timeout-minutes: 5`** — every honest run in the observed session finished in **1834 s**, so 5 min
is ~9× the slowest honest run and far under whatever reaped run 264 at 834 s. The alarm mail now
carries **`Elapsed`** (a start stamp in step 1 via `$GITHUB_ENV`; an absent stamp prints
`unknown (no start stamp)`, never a bogus number), and its "names itself in the run log" sentence is
qualified so it cannot mislead when there is no log.
## 3. The positive control
**THE UNKNOWN IS NOT CLOSED.** Whether the `if: failure()` alarm fires at all for a *reaped* job is
**still unverified**. The timeout makes the reap unreachable in practice; it does not answer what
happens inside one. Demonstrating it would mean deliberately hanging a run on `main`, which would
leave the branch red for a parallel session, so it was not done. Said in the workflow comment, the
changelog, R-265 and here — four places, none of them claiming it is answered.
```
1 BASELINE real dataset -> OK, exit 0
2 PLANT scratch copy: use.dlna missing -> walked -> CONVICTED, exit 1
"use.dlna: status 'walked' but NO evidence document cited"
3 REMOVE scratch copy deleted; committed dataset never touched
4 RE-RUN real dataset -> OK, exit 0
```
## 7. Tests
Plant → convicted → removed → clean. The gate also convicted **51 problems in my own first draft** of
the dataset (bad anchors, register ids that are not in `OPEN-ITEMS.md`, an evidence path that does not
exist) before any of this — which is the more convincing demonstration, because it was not staged.
| | before | after |
|---|---|---|
| controller | — | **1355** total (`+7` this session: 4 verdict, 3 disk-meter) |
| agent | — | **947** total (`+5` this session) |
## 4. The two known disagreements — both settled, and neither document was wrong
`go build ./... && go vet ./... && go test ./...` **green in both repos** (controller 28 packages,
agent 29), run separately from every commit. `controller_gates.py --fast` all OK; `agent_gates.py
--fast` all OK; `repo_gates.py --fast` **all 8 OK**.
**"A customer restores their own data with no help."** The map says **MISSING (as evidence)**; the
2026-08-07 walk records a customer route completed with no shell. **Not a contradiction.** The map's
row is *"A customer (**not the operator**) performs a restore via UI alone"* — it is about *who*. The
walk proves the *route*. No non-operator has ever done it, which is what the page's own neighbouring
claim already says.
The dashboard test **extracts** the meter block from the shipped template rather than copying it — a
copied block drifts and then passes while the page it covers has changed.
**The reinstall story.** The map's `PROVEN-LIVE (2026-08-04 night drill)` row is scoped in its own text
to *"a controller-data-volume rebuild — NOT a total host loss"*. The 2026-08-09 rehearsal was a
whole-host uninstall and reinstall. **The map has no row for that case at all** — a gap, not a
disagreement.
## 8. Deploy
**Would anything here have caught either one? No — and it could not have, because neither was false.**
Both are collisions of vocabulary: "customer" meaning *the route* or *a person*, "rebuild" meaning
*the guest* or *the host*. No gate detects an ambiguity that makes two true sentences look
contradictory. They surfaced only when someone tried to state them side by side. **That is the
argument for the dataset** — one id, one scope, one status — and against prose rows.
| | version | evidence |
|---|---|---|
| agent | **0.128.0** | `felhom-agent --version` on `felhom-pve`; service `active`; prior binary kept as `felhom-agent.bak-0.127.0` |
| controller | **0.210.0** | `docker ps` on guest 9201: `felhom-controller:0.210.0 Up (healthy)` |
## 5. The data file
**Endpoint-level validation of the controller UI was ATTEMPTED AND DID NOT SUCCEED — stated rather
than skipped.** What was tried: login POST against the container IP `172.17.0.2:8080` with the
mandatory `Host:` header, first with curl's cookie jar and then with the `Set-Cookie` handled
explicitly (the known `felhom_session` jar trap). Login returned `302` and `/` returned `302` back to
login both times, so the dashboard was never rendered. **Fallback observable, on the deployed
artifact rather than the source** — `grep -a` inside the running container's binary:
`SystemInfo.DiskKnown` ×1, the disk caveat line ×6, the no-verdict title ×1, `--version`
`0.210.0 (commit c732fe1)`. That proves the shipped bytes carry both fixes; it does **not** prove the
rendered page, and the template-level tests are what stand for that.
`documentation/architecture/where-felhom-stands.yaml`, 55 entries. **Every entry cites at least one
source and the gate proves it** (`check_stands.py` rule 1). YAML rather than JSON because statuses move
one line at a time and a YAML diff shows which claim moved; a JSON re-dump reflows.
## 9. The bake, and the three Day-0 values
Rules honoured: it is a **view** (every entry cites map / register / evidence); **no status was raised
in it** — the one upgrade is recorded against evidence and the map is named as the thing that must
change; and it is **regenerated, not hand-edited** for the page.
`documentation/tests/golden-0.210.0-2026-08-08/` — golden **0.210.0**, **656 787 777 B**, sha256
`b9f701fa…4c0a00`, round-trip verified, `./etc/felhom-controller-image` read **out of the downloaded
archive** → `felhom-controller:0.210.0`. Markers all green, token grep 0 with a control returning 1,
bake VM destroyed, drill disk restored to `virgin`.
## 6. The page
| field | set to | verified |
|---|---|---|
| `golden_version` | **0.210.0** | package `GET` **200**; hub dropdown offers it, `data-sha` matches the bake |
| `agent_version` | **0.128.0** | package `GET` **200**; hub dropdown offers it, `data-sha` matches the deployed binary |
| `min_agent` | **0.127.0** (unchanged) | read from the CHANGELOG header written this session |
- `where-felhom-stands.html`**generated, 54 KB, zero `<script>` tags**, same palette
(`#0b1220` / `#121b2c` / `#34d399` `#60a5fa` `#fbbf24` `#64748b`), 1600 px, A3 landscape print rules.
- Every moved status carries a visible **`changed 2026-08-09, was walked`** chip plus a *why it moved*
line — 12 of them, no diffing required.
- `Where Felhom Stands.html`**`where-felhom-stands-2026-08-09-snapshot.html`** (`git mv`, so the
space is out of every shell path), with a line in `documentation/README.md` calling it a dated
snapshot that is not maintained.
- **What the old bundle actually was**, since it aimed the fix: not merely minified — the content sat
**gzip+base64 inside a JS module map**, and the three blobs decompress to the bundler and React, with
the document itself in a JSON-escaped string on line 393. It rendered and nothing else.
**⚠ The agent was NOT published until this session checked, and it mattered.** R-221's fix is in the
**agent**; the binary had been hand-deployed and never published, so `agent_version 0.128.0` was not
selectable and a fresh install would have received 0.127.0 — the golden would have carried the
controller fixes and **not** the one the headline defect needed. Caught by checking each value was
*fetchable* instead of assuming. Published from the **live-deployed bytes**, sha-verified across the
hop first (`c6eba73b…` identical on both sides).
## 7. The corrected silence rows
**`min_agent` stays 0.127.0 deliberately:** `MinAgent` declares what the *controller* requires, and
v0.210.0 requires nothing new from the agent. R-221 is delivered by `agent_version`, not by the floor.
**R-281 is WITHDRAWN. It was wrong in both directions**, and the operator's mailbox is what settled it.
**Nothing was vouched. No hub setting was touched.** The Save is the operator's.
- **The tripwire DID fire**: `escrow_blob_served` at 10:19:41 UTC = **12:19 CEST**, eight minutes before
the verified restore.
- **False alarms fired too**: `host_down` 09:28 UTC and `node_down` 09:30 UTC, both error severity, both
`sent`, for deliberate attended work — eight operator mails in all. → **R-285**, the opposite gap
from the one filed.
## 10. Gate failures remaining
**The measurement's cause IS established.** The P7 query copied `/data/hub.db` **without
`hub.db-wal`**; the hub runs SQLite in WAL mode, so everything after the last checkpoint was invisible.
**Signature, exact:** P7 reported *"2 events all day, newest `db_dump_completed` 00:30:07"*, and the
number of rows on 08-09 at or before 00:30:07 is **exactly 2**.
**None.** `golden_currency_gate.py` went red the moment the controller was bumped — correct, and
closed by the bake, not by `--no-verify`. **No `--no-verify` anywhere in this session.**
The two obvious alternatives were **tested and refuted**, not waved away: a **timezone offset** — all
nine mailbox stamps equal the hub's UTC + 2 h exactly, so the window was right; and a **wrong key or
wrong store** — the same table and key return the correct rows now. A live re-run **cannot** reproduce
the fault because the WAL has since been checkpointed, and that is stated rather than dressed up as a
reproduction.
## 11. Register
**The lesson, filed as R-286:** the control was drawn from the *same stale snapshot* as the
measurement, so it agreed. **A control must come from a different channel.** The independent channel —
the mailbox — was available the whole time. This is also a trap `operations/nodes.md` already
documents, and which I had avoided correctly earlier in the same session.
**Closed:** R-221, R-259, R-258, R-265 (the last with its unknown explicitly still open).
**Minted:** **R-266** — the failed root `statfs` still travels to the hub as a 0-of-0 disk; ranked
low because it is the quiet direction, and now a two-repo wire change governed by G-1's gate.
**Highest ID moved R-265 → R-266.** **G-3 unblocked** in `ROADMAP.md`; **CONTEXT S-39** rules the
convention.
## 8. Register
**Still open, untouched:** R-246, R-255, R-256, R-257, R-261, R-262, R-263, R-264, R-240, R-243,
R-202, R-213, R-244, R-214/R-235, C7's test-comment half, and G-8's other half.
**Ceiling moved R-284 → R-290.** Opened: **R-285** (planned reinstall pages the operator), **R-286**
(same-channel control), **R-287** (the CI red is true), **R-288** (the capability map is unreadable),
**R-289** (R-182's row is stale), **R-290** (map rows cite no evidence). Withdrawn: **R-281**.
## 12. The capability-map row
## 9. Part 4 — and I did not change the workflow, on purpose
`00-capability-map.md:93`*"A failed per-app Tier-1 backup reaches the OPERATOR"*. **Checked, and
it was NOT claiming something untrue:** it is about the operator notification path and claims nothing
about what `/backups/apps` draws. But its narrative — *"the page you open to ask whether ONE app is
backed up"* — invites the wrong reading, and the adjacent thing WAS false: the page's tick was green
on presence until v0.210.0, so the two halves disagreed and only the operator half was true. The row
now records that.
**Which gate:** `published` / `scripts/check-published-versions.py`.
## 13. Observations — noticed, NOT acted on
**The premise is wrong in every particular. It is not ref-sensitive.** The gate enumerates releases
from the **Gitea tags API** (`main()`, `/api/v1/repos/admin/felhom-agent/tags?limit=200`), so the
checked-out ref is irrelevant — and the two previous tag pushes **passed** (run 190 `v0.126.0`, run 216
`v0.127.0`).
1. **Other collectors in `info_linux.go` return silently on error**`readLoadAvg`, `readMemInfo`
and the temperature read. Only the disk one was traced to a customer-visible surface, and the
change was deliberately not widened into a refactor of that file.
2. **The tick's recency weakness stands.** A tick over a three-week-old restore point is still a
tick. Adding an age threshold means inventing a number; the time is printed beside it.
3. **The controller UI could not be driven headlessly this session** (§8). Worth one session to
re-establish the documented headless login, because "invoke the endpoint the UI invokes" is this
project's standard validation method and it is currently unavailable for the controller.
4. **`HDDKnown` is wired but has no template consumer yet** — the HDD path renders through
`StorageBars`, which has its own `Disconnected` state. Adding the flag there is the natural next
step of the S-39 conversion and is part of G-3's survey, not this session.
**What is true:** run 267 (main, `28ba8593b8`, 08-08 14:29 UTC) printed `ok v0.120.0: binary
downloadable`. Run 284 (**the same commit**, on the tag, 08-09 09:30 UTC) printed
`FAIL v0.120.0 — HTTP 404`. A published release became uninstallable between the two. The registry now
holds exactly the ten newest versions; `0.128.0` was published **14:47 UTC, eighteen minutes after run
267**.
**Who removed `0.120.0` is NOT established, and I will not guess:** `package_cleanup_rule` is empty
(queried in Postgres), `app.ini` sets no limit, `publish-agent.sh:77` only pre-deletes the version it
is publishing, the Gitea pod has 53 days uptime and 0 restarts, and **no `DELETE` on the packages API
appears in 48 h of router logs**. The internal `[cron.cleanup_packages]` `@midnight` job falls in the
window and would leave no router line — **a leading candidate, not a conclusion.**
**So nothing was silenced and no workflow file was changed.** The red is a **true positive** — a tagged
version that cannot be installed is the exact R-115 defect the gate exists to catch, and muting it
would hide the next one. **What it therefore still does not check: nothing. Nothing was disabled.**
The honest fixes — bound the gate to versions at or above the vouched `min_agent` floor (0.127.0 today;
nothing can install 0.120.0), or retire tags whose packages go — are release decisions, and §8 forbids
fixing findings here. Filed as **R-287**.
**Also measured, and it is good news:** the failure alarm did send —
`RESEND-ACCEPTED id=fa1a7a83-714f-4357-b0ca-d3c4bb7ae73f`.
## 10. Observations — noticed, not acted on
- **`fail.app-crash` and `fault.operator-email` are the same unobserved thing** seen from two sides:
the digest is wired and correct in source, and no one has watched it arrive.
- **The verification depth is recorded per claim** (`depth: source-read | register+map |
needs-hardware`). 23 of 55 got a source or evidence read; the rest were checked against the register
and map only. That is on the face of the data rather than implied by a green tick.
- **The capability map is the thing that actually needs fixing.** The dataset now disagrees with it for
twelve rows, and the dataset is only a view — **the map still says PROVEN-LIVE for all twelve** (R-290).
- `documentation/audits/` holds 131 files and `tests/` 37; the evidence exists in quantity. The gap is
that the map does not point at it.
- **Out of scope and left alone, as instructed:** every finding above, the capability-map restructure
(R-288), and the two guards owed from yesterday.