diff --git a/REPORT.md b/REPORT.md index 69f863c..747b66b 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,184 +1,191 @@ -# REPORT — the seed that never ran twice, and three pictures that were not true (2026-08-08) +# REPORT — making the picture true (2026-08-09, unattended) -Four defects of one family, each with a source-verified mechanism, tests and red-proofs. -Agent **v0.128.0** · controller **v0.210.0** · `gates.yml` (workflow only). **The hub was not touched, -not bumped and not deployed.** +Read-only against all live infrastructure. **Both demo machines were powered off and in transit; no +box was probed, woken or waited on.** Claims that only a running box could settle are marked +`needs-hardware`, which is a verdict, not a gap. -## 1. Part 1's live sequence — RUN, OPERATOR-PRESENT, PASSED +--- -On `demo-felhom` guest 9201, agent 0.128.0. Backup taken first and verified byte-identical -(`sha256 eff18437…`), pbsdr marker sha recorded and **unchanged at every step** -(`fd4c97b5…`). +## 1. The verdict table -**Step 3 — the preflight with the key removed (today's defect, live):** +**55 claims. Statuses moved on 12 of them — all downwards.** `walked 32 → 20`, `built 5 → 17`; +`partial 14` and `missing 4` unchanged. Full per-claim detail with sources is in +`documentation/architecture/where-felhom-stands.yaml`. -``` -{"id": "pbs_storage_id", "ok": false, "detail": "escrow.pbs_storage_id not configured"} -``` - -**Step 4 — one 60 s tick later, same daemon:** - -``` -{"id": "pbs_storage_id", "ok": true, "detail": "felhom-pbs"} -``` - -**No restart:** agent MainPID **1993397 before and 1993397 after** — the same process. The wizard -would have refused at 18:04:16Z and would pass at 18:05:44Z, with nobody touching the box. - -**Step 5 — nothing else moved:** 45 keys in the backup, 45 now, **added none / removed none / -changed none**. The one key is back with its original value. - -**Positive control:** the preflight was exercised BEFORE the change and reported `ok: true` on the -same row, so a green afterwards is a measurement and not an artefact of the probe. - -## 2. The writer of `agent.json` — ESTABLISHED - -`step_agent_config`, **`felhom.eu/scripts/felhom-host-install.sh:2396`**; the Python render at -**`:2449`**; the `O_TRUNC` write at **`:2579`**. `PRESERVE_FROM` defaults empty (**`:256`**) and is set -only by an explicit `--preserve-from` (**`:1246`**). **The render never writes an `escrow` section at -all** — grep over the whole heredoc: zero hits. The pbsdr marker is host-side -(`/pbsdr/marker.json`) and survives. **R-221's attribution was correct.** - -A rebuild is only the case that was *measured*; the same hole opens for a hand-edited or restored -config, which is the honest reason the fix is at the seam rather than in the installer. - -## 3. Red-proofs — 6 of 6, each with the mutation asserted applied - -| # | mutation | assertion it applied | outcome | +| claim | was | now | why | |---|---|---|---| -| 1 | **Part 1 / Scenario A: remove the new seed call** | marker `MUTATED: the R-221 re-assert removed` present | **RED** — and **yes, it failed against today's tree**, with the intended message ✔ | -| 2 | Part 1: remove the early return as well | marker `MUTATED: early return deleted` present | **RED** — the zero-Proxmox-calls assertion is load-bearing, not decorative ✔ | -| 3 | Part 3 / F: revert to `status.LastDBDump.Success` | marker present | **RED** — app X's false green returns ✔ | -| 4 | Part 3 / G: map "no result" to `ok` | marker present | **RED** — green-on-presence returns ✔ | -| 5 | Part 2 / D: ignore `DiskKnown` in the template | marker present | **RED** — „0.0 GB / 0.0 GB (0%)" in the nominal colour returns ✔ | -| 6 | Part 2 / E: force the flag false | marker present | **RED** — a healthy box is shown losing its numbers ✔ | +| `install.installer-by-tag` | walked | **built** | gate 6 asserts the manifest names an installer tag; no walk of a rollback on file | +| `use.lifecycle` | walked | **built** | no walk document cited | +| `drives.enrol` | walked | **built** | the 08-09 walk exercised RE-attach (which failed, R-280); first enrolment of a NEW drive has no walk | +| `drives.migrate` | walked | **built** | no walk document cited | +| `backup.tier1` | walked | **built** | no walk document cited | +| `backup.whole-machine` | walked | **built** | no walk document cited | +| `backup.restore-proof` | walked | **built** | no walk cited, **and the last recorded restore-test on demo-hp FAILED** (2026-08-05) | +| `fault.selfheal` | walked | **built** | no walk document cited | +| `fault.operator-email` | walked | **built** | source-verified as correct, but no run observed delivering | +| `fail.drive-filling` | walked | **built** | no walk document cited | +| `fail.lost-recovery-code` | walked | **built** | by-design refusal; no walk document cited | +| `fail.hub-down` | walked | **built** | no walk document cited | -All six restored and re-verified green. **Answer to the question asked directly: the Part 1 test DID -fail against today's tree.** +**Upgraded: 1.** `install.byo` — the page said *"the first real one has not happened"*. A real +`--mode byo` install completed on demo-hp on 2026-08-09 (`Day-0 provision SUCCESS`, 3 m 49 s). Still +not a customer's own hardware, so not *walked*, but the sentence was false. -## 4. The Hungarian strings as shipped +**`needs-hardware`: 4** — `use.lan-fallback`, `backup.restore-proof`, `fail.disk-failing`, +`fail.internet-down`. Each needs an observation on a running box; each says which. -- `„A tárhely mérete most nem olvasható ki."` — the disk caveat line -- `„nem ismert"` — the short label in the value slot -- `„Erről a mentésről nincs eredményünk."` — the `title` on the no-verdict backup mark +**Confirmed: 38.** Seven of those were re-confirmed against live source or the live hub tonight rather +than against paperwork: the tripwire, the off-site repository, the claim path, the catalogue, the +tunnel, the reset code and the operator-email digest. -## 5. The §7.3 truth table as implemented +## 2. Every downgrade, with the coupling that broke -| this app's own most recent dump result | restore point | verdict | -|---|---|---| -| any of its databases failed | yes | `error` | -| all clean | yes | `ok` | -| none recorded | yes | **no icon**, time only, with the title above | -| any | no | no tier-1 row, unchanged | +The rule is *"a proof is about the code that existed when it ran"*. **It did not fire the way the task +expected.** Not one downgrade came from code moving under an old proof. **All twelve came from step 1 +of the same rule — the cited evidence does not exist.** -**Recency was left alone**, deliberately: an age threshold means inventing a number, and the time is -already printed beside the icon. Recorded as an observation. +Measured: of the 28 capability-map rows behind the page's claims, **8 carry a `tests/` or `audits/` +path in their evidence column and 20 carry prose only.** The green dots were being drawn from rows +that cite an argument, not a walk. Filed as **R-290**. -## 6. The CI timeout, and what is still unknown +**And the decay ran the other way once.** `fault.operator-email` — *"one mail per run, every failing +app named"* — I first took to be contradicted by R-182 (open, *"tells the operator about ONE app and +silently swallows every other"*). Reading live source: the digest `backup_run_failures` is allowlisted +(`hub/internal/api/handler.go:1837`), operator-only (`notify/dispatcher.go:423`) and templated +(`notify/templates.go:48`); `recovery_unit_capture_failed` is record-only (`dispatcher.go:376`); a +cooldown drop now logs a `suppressed` row (`dispatcher.go:314-330`). **The claim is right and the +register row is stale** — filed as **R-289**. The session went looking for stale proofs and found a +stale defect. -**`timeout-minutes: 5`** — every honest run in the observed session finished in **18–34 s**, so 5 min -is ~9× the slowest honest run and far under whatever reaped run 264 at 834 s. The alarm mail now -carries **`Elapsed`** (a start stamp in step 1 via `$GITHUB_ENV`; an absent stamp prints -`unknown (no start stamp)`, never a bogus number), and its "names itself in the run log" sentence is -qualified so it cannot mislead when there is no log. +## 3. The positive control -**THE UNKNOWN IS NOT CLOSED.** Whether the `if: failure()` alarm fires at all for a *reaped* job is -**still unverified**. The timeout makes the reap unreachable in practice; it does not answer what -happens inside one. Demonstrating it would mean deliberately hanging a run on `main`, which would -leave the branch red for a parallel session, so it was not done. Said in the workflow comment, the -changelog, R-265 and here — four places, none of them claiming it is answered. +``` +1 BASELINE real dataset -> OK, exit 0 +2 PLANT scratch copy: use.dlna missing -> walked -> CONVICTED, exit 1 + "use.dlna: status 'walked' but NO evidence document cited" +3 REMOVE scratch copy deleted; committed dataset never touched +4 RE-RUN real dataset -> OK, exit 0 +``` -## 7. Tests +Plant → convicted → removed → clean. The gate also convicted **51 problems in my own first draft** of +the dataset (bad anchors, register ids that are not in `OPEN-ITEMS.md`, an evidence path that does not +exist) before any of this — which is the more convincing demonstration, because it was not staged. -| | before | after | -|---|---|---| -| controller | — | **1355** total (`+7` this session: 4 verdict, 3 disk-meter) | -| agent | — | **947** total (`+5` this session) | +## 4. The two known disagreements — both settled, and neither document was wrong -`go build ./... && go vet ./... && go test ./...` **green in both repos** (controller 28 packages, -agent 29), run separately from every commit. `controller_gates.py --fast` all OK; `agent_gates.py ---fast` all OK; `repo_gates.py --fast` **all 8 OK**. +**"A customer restores their own data with no help."** The map says **MISSING (as evidence)**; the +2026-08-07 walk records a customer route completed with no shell. **Not a contradiction.** The map's +row is *"A customer (**not the operator**) performs a restore via UI alone"* — it is about *who*. The +walk proves the *route*. No non-operator has ever done it, which is what the page's own neighbouring +claim already says. -The dashboard test **extracts** the meter block from the shipped template rather than copying it — a -copied block drifts and then passes while the page it covers has changed. +**The reinstall story.** The map's `PROVEN-LIVE (2026-08-04 night drill)` row is scoped in its own text +to *"a controller-data-volume rebuild — NOT a total host loss"*. The 2026-08-09 rehearsal was a +whole-host uninstall and reinstall. **The map has no row for that case at all** — a gap, not a +disagreement. -## 8. Deploy +**Would anything here have caught either one? No — and it could not have, because neither was false.** +Both are collisions of vocabulary: "customer" meaning *the route* or *a person*, "rebuild" meaning +*the guest* or *the host*. No gate detects an ambiguity that makes two true sentences look +contradictory. They surfaced only when someone tried to state them side by side. **That is the +argument for the dataset** — one id, one scope, one status — and against prose rows. -| | version | evidence | -|---|---|---| -| agent | **0.128.0** | `felhom-agent --version` on `felhom-pve`; service `active`; prior binary kept as `felhom-agent.bak-0.127.0` | -| controller | **0.210.0** | `docker ps` on guest 9201: `felhom-controller:0.210.0 Up (healthy)` | +## 5. The data file -**Endpoint-level validation of the controller UI was ATTEMPTED AND DID NOT SUCCEED — stated rather -than skipped.** What was tried: login POST against the container IP `172.17.0.2:8080` with the -mandatory `Host:` header, first with curl's cookie jar and then with the `Set-Cookie` handled -explicitly (the known `felhom_session` jar trap). Login returned `302` and `/` returned `302` back to -login both times, so the dashboard was never rendered. **Fallback observable, on the deployed -artifact rather than the source** — `grep -a` inside the running container's binary: -`SystemInfo.DiskKnown` ×1, the disk caveat line ×6, the no-verdict title ×1, `--version` → -`0.210.0 (commit c732fe1)`. That proves the shipped bytes carry both fixes; it does **not** prove the -rendered page, and the template-level tests are what stand for that. +`documentation/architecture/where-felhom-stands.yaml`, 55 entries. **Every entry cites at least one +source and the gate proves it** (`check_stands.py` rule 1). YAML rather than JSON because statuses move +one line at a time and a YAML diff shows which claim moved; a JSON re-dump reflows. -## 9. The bake, and the three Day-0 values +Rules honoured: it is a **view** (every entry cites map / register / evidence); **no status was raised +in it** — the one upgrade is recorded against evidence and the map is named as the thing that must +change; and it is **regenerated, not hand-edited** for the page. -`documentation/tests/golden-0.210.0-2026-08-08/` — golden **0.210.0**, **656 787 777 B**, sha256 -`b9f701fa…4c0a00`, round-trip verified, `./etc/felhom-controller-image` read **out of the downloaded -archive** → `felhom-controller:0.210.0`. Markers all green, token grep 0 with a control returning 1, -bake VM destroyed, drill disk restored to `virgin`. +## 6. The page -| field | set to | verified | -|---|---|---| -| `golden_version` | **0.210.0** | package `GET` **200**; hub dropdown offers it, `data-sha` matches the bake | -| `agent_version` | **0.128.0** | package `GET` **200**; hub dropdown offers it, `data-sha` matches the deployed binary | -| `min_agent` | **0.127.0** (unchanged) | read from the CHANGELOG header written this session | +- `where-felhom-stands.html` — **generated, 54 KB, zero `