Files
felhom.eu/REPORT.md
T
admin 6088afcbed
gates / gates (push) Successful in 21s
Verify the standing picture against source: 12 downgrades, and the decay ran both ways
55 claims verified. Twelve moved, all downwards: walked 32 -> 20, built 5 -> 17.
Register ceiling R-284 -> R-290.

THE RULE DID NOT FIRE THE WAY IT WAS EXPECTED TO. Not one downgrade came from
code moving under an old proof. All twelve came from step 1 of the same rule --
the cited evidence does not exist. Measured: of the 28 capability-map rows
behind the page's claims, 8 carry a tests/ or audits/ path and 20 carry prose
only. The green dots were drawn from rows that cite an argument, not a walk
(R-290). The map, not the dataset, is what needs fixing -- it still says
PROVEN-LIVE for all twelve.

And once it ran backwards: fault.operator-email looked contradicted by R-182,
but live source shows the backup_run_failures digest allowlisted, operator-only
and templated, with recovery_unit_capture_failed now record-only. The claim is
right and the REGISTER ROW is stale (R-289). The session went looking for stale
proofs and found a stale defect.

R-281 WITHDRAWN -- wrong in both directions, settled by the operator's mailbox.
The tripwire DID fire (escrow_blob_served 10:19:41Z = 12:19 CEST) and false
error-severity alarms fired too, for deliberate attended work (R-285). The
measurement's cause is ESTABLISHED: the P7 query copied hub.db without hub.db-wal,
and the signature is exact -- it reported "2 events all day, newest 00:30:07",
and the rows at or before 00:30:07 number exactly 2. Timezone and wrong-key were
tested and refuted. The control had been drawn from the same stale snapshot as
the measurement, which is why it agreed (R-286).

Part 4: NO WORKFLOW CHANGED, deliberately. The gate is not ref-sensitive -- it
enumerates from the Gitea tags API, and both previous tag pushes passed. The red
is TRUE: run 267 saw v0.120.0 downloadable, run 284 on the same commit saw 404.
Who deleted the package is NOT established and is not guessed (R-287).

The page is now generated from where-felhom-stands.yaml by scripts/render_stands.py:
static, zero script tags, every moved status carrying a visible "changed, was X"
chip. The React bundle -- whose content was gzip+base64 inside a JS module map --
is kept as a dated snapshot. scripts/check_stands.py gates the data and convicted
51 problems in my own first draft before the staged positive control ever ran.
2026-08-09 18:40:49 +02:00

12 KiB

REPORT — making the picture true (2026-08-09, unattended)

Read-only against all live infrastructure. Both demo machines were powered off and in transit; no box was probed, woken or waited on. Claims that only a running box could settle are marked needs-hardware, which is a verdict, not a gap.


1. The verdict table

55 claims. Statuses moved on 12 of them — all downwards. walked 32 → 20, built 5 → 17; partial 14 and missing 4 unchanged. Full per-claim detail with sources is in documentation/architecture/where-felhom-stands.yaml.

claim was now why
install.installer-by-tag walked built gate 6 asserts the manifest names an installer tag; no walk of a rollback on file
use.lifecycle walked built no walk document cited
drives.enrol walked built the 08-09 walk exercised RE-attach (which failed, R-280); first enrolment of a NEW drive has no walk
drives.migrate walked built no walk document cited
backup.tier1 walked built no walk document cited
backup.whole-machine walked built no walk document cited
backup.restore-proof walked built no walk cited, and the last recorded restore-test on demo-hp FAILED (2026-08-05)
fault.selfheal walked built no walk document cited
fault.operator-email walked built source-verified as correct, but no run observed delivering
fail.drive-filling walked built no walk document cited
fail.lost-recovery-code walked built by-design refusal; no walk document cited
fail.hub-down walked built no walk document cited

Upgraded: 1. install.byo — the page said "the first real one has not happened". A real --mode byo install completed on demo-hp on 2026-08-09 (Day-0 provision SUCCESS, 3 m 49 s). Still not a customer's own hardware, so not walked, but the sentence was false.

needs-hardware: 4use.lan-fallback, backup.restore-proof, fail.disk-failing, fail.internet-down. Each needs an observation on a running box; each says which.

Confirmed: 38. Seven of those were re-confirmed against live source or the live hub tonight rather than against paperwork: the tripwire, the off-site repository, the claim path, the catalogue, the tunnel, the reset code and the operator-email digest.

2. Every downgrade, with the coupling that broke

The rule is "a proof is about the code that existed when it ran". It did not fire the way the task expected. Not one downgrade came from code moving under an old proof. All twelve came from step 1 of the same rule — the cited evidence does not exist.

Measured: of the 28 capability-map rows behind the page's claims, 8 carry a tests/ or audits/ path in their evidence column and 20 carry prose only. The green dots were being drawn from rows that cite an argument, not a walk. Filed as R-290.

And the decay ran the other way once. fault.operator-email"one mail per run, every failing app named" — I first took to be contradicted by R-182 (open, "tells the operator about ONE app and silently swallows every other"). Reading live source: the digest backup_run_failures is allowlisted (hub/internal/api/handler.go:1837), operator-only (notify/dispatcher.go:423) and templated (notify/templates.go:48); recovery_unit_capture_failed is record-only (dispatcher.go:376); a cooldown drop now logs a suppressed row (dispatcher.go:314-330). The claim is right and the register row is stale — filed as R-289. The session went looking for stale proofs and found a stale defect.

3. The positive control

1 BASELINE  real dataset                                        -> OK,        exit 0
2 PLANT     scratch copy: use.dlna  missing -> walked           -> CONVICTED, exit 1
            "use.dlna: status 'walked' but NO evidence document cited"
3 REMOVE    scratch copy deleted; committed dataset never touched
4 RE-RUN    real dataset                                        -> OK,        exit 0

Plant → convicted → removed → clean. The gate also convicted 51 problems in my own first draft of the dataset (bad anchors, register ids that are not in OPEN-ITEMS.md, an evidence path that does not exist) before any of this — which is the more convincing demonstration, because it was not staged.

4. The two known disagreements — both settled, and neither document was wrong

"A customer restores their own data with no help." The map says MISSING (as evidence); the 2026-08-07 walk records a customer route completed with no shell. Not a contradiction. The map's row is "A customer (not the operator) performs a restore via UI alone" — it is about who. The walk proves the route. No non-operator has ever done it, which is what the page's own neighbouring claim already says.

The reinstall story. The map's PROVEN-LIVE (2026-08-04 night drill) row is scoped in its own text to "a controller-data-volume rebuild — NOT a total host loss". The 2026-08-09 rehearsal was a whole-host uninstall and reinstall. The map has no row for that case at all — a gap, not a disagreement.

Would anything here have caught either one? No — and it could not have, because neither was false. Both are collisions of vocabulary: "customer" meaning the route or a person, "rebuild" meaning the guest or the host. No gate detects an ambiguity that makes two true sentences look contradictory. They surfaced only when someone tried to state them side by side. That is the argument for the dataset — one id, one scope, one status — and against prose rows.

5. The data file

documentation/architecture/where-felhom-stands.yaml, 55 entries. Every entry cites at least one source and the gate proves it (check_stands.py rule 1). YAML rather than JSON because statuses move one line at a time and a YAML diff shows which claim moved; a JSON re-dump reflows.

Rules honoured: it is a view (every entry cites map / register / evidence); no status was raised in it — the one upgrade is recorded against evidence and the map is named as the thing that must change; and it is regenerated, not hand-edited for the page.

6. The page

  • where-felhom-stands.htmlgenerated, 54 KB, zero <script> tags, same palette (#0b1220 / #121b2c / #34d399 #60a5fa #fbbf24 #64748b), 1600 px, A3 landscape print rules.
  • Every moved status carries a visible changed 2026-08-09, was walked chip plus a why it moved line — 12 of them, no diffing required.
  • Where Felhom Stands.htmlwhere-felhom-stands-2026-08-09-snapshot.html (git mv, so the space is out of every shell path), with a line in documentation/README.md calling it a dated snapshot that is not maintained.
  • What the old bundle actually was, since it aimed the fix: not merely minified — the content sat gzip+base64 inside a JS module map, and the three blobs decompress to the bundler and React, with the document itself in a JSON-escaped string on line 393. It rendered and nothing else.

7. The corrected silence rows

R-281 is WITHDRAWN. It was wrong in both directions, and the operator's mailbox is what settled it.

  • The tripwire DID fire: escrow_blob_served at 10:19:41 UTC = 12:19 CEST, eight minutes before the verified restore.
  • False alarms fired too: host_down 09:28 UTC and node_down 09:30 UTC, both error severity, both sent, for deliberate attended work — eight operator mails in all. → R-285, the opposite gap from the one filed.

The measurement's cause IS established. The P7 query copied /data/hub.db without hub.db-wal; the hub runs SQLite in WAL mode, so everything after the last checkpoint was invisible. Signature, exact: P7 reported "2 events all day, newest db_dump_completed 00:30:07", and the number of rows on 08-09 at or before 00:30:07 is exactly 2.

The two obvious alternatives were tested and refuted, not waved away: a timezone offset — all nine mailbox stamps equal the hub's UTC + 2 h exactly, so the window was right; and a wrong key or wrong store — the same table and key return the correct rows now. A live re-run cannot reproduce the fault because the WAL has since been checkpointed, and that is stated rather than dressed up as a reproduction.

The lesson, filed as R-286: the control was drawn from the same stale snapshot as the measurement, so it agreed. A control must come from a different channel. The independent channel — the mailbox — was available the whole time. This is also a trap operations/nodes.md already documents, and which I had avoided correctly earlier in the same session.

8. Register

Ceiling moved R-284 → R-290. Opened: R-285 (planned reinstall pages the operator), R-286 (same-channel control), R-287 (the CI red is true), R-288 (the capability map is unreadable), R-289 (R-182's row is stale), R-290 (map rows cite no evidence). Withdrawn: R-281.

9. Part 4 — and I did not change the workflow, on purpose

Which gate: published / scripts/check-published-versions.py.

The premise is wrong in every particular. It is not ref-sensitive. The gate enumerates releases from the Gitea tags API (main(), /api/v1/repos/admin/felhom-agent/tags?limit=200), so the checked-out ref is irrelevant — and the two previous tag pushes passed (run 190 v0.126.0, run 216 v0.127.0).

What is true: run 267 (main, 28ba8593b8, 08-08 14:29 UTC) printed ok v0.120.0: binary downloadable. Run 284 (the same commit, on the tag, 08-09 09:30 UTC) printed FAIL v0.120.0 — HTTP 404. A published release became uninstallable between the two. The registry now holds exactly the ten newest versions; 0.128.0 was published 14:47 UTC, eighteen minutes after run 267.

Who removed 0.120.0 is NOT established, and I will not guess: package_cleanup_rule is empty (queried in Postgres), app.ini sets no limit, publish-agent.sh:77 only pre-deletes the version it is publishing, the Gitea pod has 53 days uptime and 0 restarts, and no DELETE on the packages API appears in 48 h of router logs. The internal [cron.cleanup_packages] @midnight job falls in the window and would leave no router line — a leading candidate, not a conclusion.

So nothing was silenced and no workflow file was changed. The red is a true positive — a tagged version that cannot be installed is the exact R-115 defect the gate exists to catch, and muting it would hide the next one. What it therefore still does not check: nothing. Nothing was disabled. The honest fixes — bound the gate to versions at or above the vouched min_agent floor (0.127.0 today; nothing can install 0.120.0), or retire tags whose packages go — are release decisions, and §8 forbids fixing findings here. Filed as R-287.

Also measured, and it is good news: the failure alarm did send — RESEND-ACCEPTED id=fa1a7a83-714f-4357-b0ca-d3c4bb7ae73f.

10. Observations — noticed, not acted on

  • fail.app-crash and fault.operator-email are the same unobserved thing seen from two sides: the digest is wired and correct in source, and no one has watched it arrive.
  • The verification depth is recorded per claim (depth: source-read | register+map | needs-hardware). 23 of 55 got a source or evidence read; the rest were checked against the register and map only. That is on the face of the data rather than implied by a green tick.
  • The capability map is the thing that actually needs fixing. The dataset now disagrees with it for twelve rows, and the dataset is only a view — the map still says PROVEN-LIVE for all twelve (R-290).
  • documentation/audits/ holds 131 files and tests/ 37; the evidence exists in quantity. The gap is that the map does not point at it.
  • Out of scope and left alone, as instructed: every finding above, the capability-map restructure (R-288), and the two guards owed from yesterday.