where-felhom-stands: bring the picture up to 2026-08-22, and stop the page disagreeing with its source
gates / gates (push) Successful in 15s

Eight claims re-checked against the drill and the v0.218.0 fixes; three moved, all downward.

  backup.offsite            walked -> partial. "18 snapshots, daily, unbroken" was true on
                            2026-08-09 and false by 2026-08-21: the next snapshot after that date
                            was put there by hand, twelve days later. The rebuild lost the target
                            and the per-app switches came back off, so a run reported "backup OK:
                            0 app(s) backed up".
  fail.wiped-reinstalled.data  walked -> partial. A real reinstall orphans BOTH off-premises tiers:
                            restic silently for 12 days (R-193), and the PBS archives from before
                            the reinstall cannot be opened by the rebuilt box at all (R-366).
  backup.fill-warning       walked -> partial. The warning fires correctly, but the watcher runs
                            once a day, so a filesystem that fills at 03:31 goes unannounced for
                            ~24 h. Watched silent while a volume sat at 99%.

Five re-checked and held: backup.tier1 and recover.byte-identical carry the R-355/R-354 story and
their fixes; backup.restore-proof stays grey for a sharper reason (orphaned archives, not an
untested tier); backup.sikeres gains two fresh instances; fail.customer-self-restore records that
R-356 now blocks 40 of 53 apps regardless of who is driving.

render_stands.py: the header's commit shas were hardcoded, so the page cited the August 9th commits
while the YAML said otherwise - the stale-build-product failure the renderer exists to prevent. They
are parsed now. The count beside them said "15 status(es) moved in that pass" when 15 was every
recorded move ever; it now separates the two numbers.

check_stands passes, and was itself proven able to convict first: a claim marked `missing` flipped to
`walked` in a scratch copy fired rule 5 by name (use.dlna), and the real file still passes.
This commit is contained in:
2026-08-22 10:36:42 +02:00
parent 877fcd2a38
commit ca543b8f69
4 changed files with 98 additions and 38 deletions
File diff suppressed because one or more lines are too long
@@ -12,11 +12,11 @@
# verdict vocabulary: confirmed | downgraded | upgraded | contested | needs-hardware
# depth: source-read (opened live source or evidence) | register+map (checked against the
# register and capability map only) | needs-hardware (cannot be settled off-box)
verified_on: 2026-08-09
verified_on: 2026-08-22
verified_against:
felhom-agent: 28ba8593b8
felhom-controller: c732fe1283
hub: 56f8aa611c
felhom-agent: 40d857b527
felhom-controller: 2da259af38
hub: 877fcd2a38
claims:
- id: install.iso-selfregister
band: journey
@@ -301,17 +301,20 @@ claims:
stage: 5
title: "App data on the machine, nightly database dumps, a copy on a second drive"
status: walked
note: "True for every app but one, and that one was found on 2026-08-21: paperless-ngx's 72-table PostgreSQL was dumped nightly and written into a folder named after an app that does not exist, so it never entered the recovery unit, the second-drive copy or the off-site copy. Fixed and proven live 2026-08-22 (R-355) — the dump is now in the app's own unit and in the off-site snapshot for the first time. A catalogue-wide sweep, itself proven able to convict a planted case, says 1 of 53 was affected."
sources:
- capability-map: "Tier-2 secondary-drive copy: class-driven legs"
- register: "R-355"
- evidence: "audits/CAMPAIGN-8-backup-restore-2026-07-27.md"
- evidence: "audits/DRILL-backup-truth-2026-08-21/evidence/phase4-paperless/FINDING.txt"
changed:
from: built
also_moved: 2026-08-09 walked -> built
reason_superseded: "no walk document cited"
reason: "receipt found 2026-08-10: an adversarial, destructive, unattended overnight campaign across both boxes and ep0, with A2 (one quiesce, two tiers) proven end to end. Map already read PROVEN-LIVE."
verified:
date: 2026-08-09
verdict: upgraded
date: 2026-08-22
verdict: confirmed
depth: source-read
- id: backup.whole-machine
band: journey
@@ -332,59 +335,75 @@ claims:
band: journey
stage: 5
title: "An encrypted off-site copy, sealed with a key the operator cannot read"
status: walked
note: "Verified tonight from the repo itself: 18 snapshots, daily, unbroken."
status: partial
note: "The COPY is real and the key is still unreadable to us — what failed is DAILY and UNBROKEN. The 2026-08-09 note said '18 snapshots, daily, unbroken'; the next snapshot after 2026-08-09 08:30 was 2026-08-21 22:17, put there by hand during the drill. Twelve days, no alarm. Two causes in series: the 2026-08-21 rebuild lost the off-box target (R-193's shape), and after the self-heal restored it EVERY per-app switch was still off, so the first run logged 'backup OK: 0 app(s) backed up, 14s'."
sources:
- capability-map: "Offsite (restic → Hetzner Storage Box)"
- register: "R-199"
- register: "R-193"
- evidence: "audits/REHEARSAL-byo-reinstall-2026-08-09.md"
- evidence: "audits/REPORT-DRILL-backup-truth-2026-08-21.md"
changed:
from: walked
reason: "the claim is about a CONTINUING daily copy; a 12-day silent gap was found on 2026-08-21 and nothing reported it"
verified:
date: 2026-08-09
verdict: confirmed
date: 2026-08-22
verdict: downgraded
depth: source-read
- id: backup.restore-proof
band: journey
stage: 5
title: "The backups prove themselves: a restore is actually performed, unattended, on every tier, on both machines"
status: built
note: "'on both machines, unattended, every tier' is a continuing claim about scheduled runs. Both boxes are off; the last recorded restore-test on demo-hp FAILED (notification_log 2026-08-05 restore_test_failed). Cannot be settled tonight."
note: "STILL GREY, and for a sharper reason than in August. The scheduler runs and the check works — it failed again on 2026-08-21, unprompted, and named the cause exactly: after the reinstall the box presents a different PBS key than its own older archives were sealed with, so those archives cannot be opened at all (R-366). A tier whose archives are orphaned is a different alarm from a tier whose test failed, and only the second is being said."
sources:
- capability-map: "Restore-proof is UNATTENDED — the scheduler covers EVERY tier"
- register: "R-86"
- register: "R-366"
- evidence: "audits/REPORT-DRILL-backup-truth-2026-08-21.md"
changed:
from: walked
reason: "no walk document cited, and the last recorded restore-test on demo-hp FAILED (2026-08-05)"
decay: "PROOF-DECAY RULE FIRED (first time it has). A receipt EXISTS - architecture/_recovery-inventory-2026-07-28.md carries live journal lines for scheduled restore-tests on both boxes and both tiers - but it is superseded by later observation: demo-hp logged restore_test_failed on 2026-08-05, and the box has since been wiped and reinstalled (2026-08-09). The claim is about a CONTINUING scheduled behaviour, so a 2026-07-28 observation cannot carry it. Stays grey until a scheduled restore-test is seen passing on the rebuilt box. THE CAPABILITY MAP STILL READS PROVEN-LIVE (2026-08-03) AND IS NOW THE THING OUT OF STEP."
worse_2026_08_22: "It failed AGAIN, unprompted, on 2026-08-21 21:59 - and the cause is worse than 'untested'. Hub event 3016: the PBS archive of 2026-08-18 could not be restored because the manifest's key does not match the key the rebuilt box now presents. The archives that predate the 2026-08-21 reinstall are UNREADABLE to the machine that made them (R-366). Credit where due: the mechanism caught it and named the key mismatch precisely. The gap is that it is reported as 'a restore test failed' rather than 'your older whole-machine backups cannot be opened on this box'."
verified:
date: 2026-08-09
date: 2026-08-22
verdict: downgraded
depth: needs-hardware
- id: backup.fill-warning
band: journey
stage: 5
title: "The customer is warned before a drive fills, per drive, in their own language"
status: walked
note: "R-177 (no operator-triggerable run) limits testing, not the capability."
status: partial
note: "The warning is real and was SEEN firing on 2026-08-21 with the right Hungarian copy, naming the drive and the free space. What 'BEFORE' cannot survive is the cadence: the watcher runs once a day at 03:30 plus once at startup (R-363), so a filesystem that fills at 03:31 goes unannounced for ~24 h. Watched live: the 69 GB volume carrying all 40-class app data was filled to 99% and the watcher said nothing, while the backup reserve was already refusing an app per run and telling the hub about it."
sources:
- capability-map: "The customer is warned BEFORE a filesystem fills"
- register: "R-167"
- register: "R-363"
- evidence: "audits/SPIKE-r165-mp1-merge-2026-08-02.md"
- evidence: "audits/REPORT-DRILL-backup-truth-2026-08-21.md"
changed:
from: walked
reason: "a daily check cannot carry the word BEFORE; observed silent for the whole window a filesystem sat at 99%"
verified:
date: 2026-08-09
verdict: confirmed
depth: register+map
date: 2026-08-22
verdict: downgraded
depth: source-read
- id: backup.sikeres
band: journey
stage: 5
title: "A backup that covered nothing still calls itself successful"
status: partial
note: "Warning card."
note: "Warning card — and the drill found two more of it, both live. (1) An off-site run with no app selected logs 'backup OK: 0 app(s) backed up'; the card does say 'nincs kijelölt alkalmazás' beside the green tick, so this one is honest if you read past the tick. (2) A restore that placed nothing reported success: '0 fájl visszaállítva', ok=true. The second is the one that matters and it is R-353, still open. The volume half of it is fixed (R-354): the message now names what came back."
sources:
- register: "R-240"
- register: "R-353"
- register: "R-354"
- evidence: "audits/DRILL-backup-truth-2026-08-21/evidence/phase3-experiment/messages-verbatim.txt"
verified:
date: 2026-08-09
date: 2026-08-22
verdict: confirmed
depth: register+map
depth: source-read
- id: fault.selfheal
band: journey
stage: 6
@@ -460,13 +479,16 @@ claims:
stage: 7
title: "The customer's code opens the sealed package and the data returns byte for byte — including accented Hungarian filenames, verified as raw bytes"
status: walked
note: "Reproduced 2026-08-09: 4/4 byte-identical, name bytes NFC-preserved, out of snapshot 41c830db."
note: "Still true, and re-proven 2026-08-22 (5/5 byte-identical, both accented names as raw bytes) — but the SCOPE is narrower than the sentence sounds and was silently narrower still until v0.218.0. The drill found the off-site restore had NO named-volume leg at all: the archive sat in the unit, the snapshot and the checking folder and was never replayed, under a success message (R-354, fixed and proven 2026-08-22). AND 40 of the 53 catalogue apps STILL cannot run this route at all — it refuses first, saying a running app is not installed (R-356, open). So: proven for an app that declares a data drive; unproven and currently unreachable for the class whose entire dataset is a named volume."
sources:
- register: "R-201"
- register: "R-354"
- register: "R-356"
- evidence: "tests/walk5-r201-2026-08-07/journal.md"
- evidence: "audits/REHEARSAL-byo-reinstall-2026-08-09.md"
- evidence: "audits/REPORT-DRILL-backup-truth-2026-08-21.md"
verified:
date: 2026-08-09
date: 2026-08-22
verdict: confirmed
depth: source-read
- id: recover.no-shell
@@ -581,13 +603,19 @@ claims:
- id: fail.wiped-reinstalled.data
band: failures
title: "The whole machine is wiped and reinstalled — the data comes back"
status: walked
note: "4/4 byte-identical 2026-08-09."
status: partial
note: "The 2026-08-09 rehearsal really did return 4/4 byte-identical, and that stands. What the next real reinstall showed (2026-08-21, demo-hp) is that the rebuild ORPHANS BOTH OFF-PREMISES TIERS at once, quietly: the restic target was lost and needed a self-heal plus a per-app re-enable before any copy resumed (R-193), and the PBS archives from before the reinstall cannot be opened by the rebuilt box at all, because it now presents a different key (R-366). The data came back in the rehearsal because the rehearsal restored it immediately; a machine left alone after a reinstall is not protected in the meantime and nothing says so."
sources:
- register: "R-193"
- register: "R-366"
- evidence: "audits/REHEARSAL-byo-reinstall-2026-08-09.md"
- evidence: "audits/REPORT-DRILL-backup-truth-2026-08-21.md"
changed:
from: walked
reason: "a real reinstall on 2026-08-21 left both off-premises tiers broken - one silently for 12 days, the other unreadable - so 'the data comes back' holds only if someone restores it at once"
verified:
date: 2026-08-09
verdict: confirmed
date: 2026-08-22
verdict: downgraded
depth: source-read
- id: fail.wiped-reinstalled.journey
band: failures
@@ -727,11 +755,13 @@ claims:
band: failures
title: "A customer restores their own data with no help"
status: partial
note: "CONTESTED-RESOLVED: the map's MISSING is about a NON-OPERATOR performing it; the walks prove the route, not the person. Neither record was wrong."
note: "CONTESTED-RESOLVED: the map's MISSING is about a NON-OPERATOR performing it; the walks prove the route, not the person. Neither record was wrong. SINCE 2026-08-21 there is a second, harder blocker and it is not about the person at all: for the 40 of 53 apps that declare no data drive the off-site restore REFUSES before it starts, telling the customer a running app 'nincs telepítve' and to reinstall it to the same place — which those apps give them no way to choose (R-356). Those are exactly the apps whose whole dataset is a named volume. Until that is fixed, most customers cannot self-restore off-site even in principle."
sources:
- capability-map: "A customer (not the operator) performs a restore via UI alone"
- register: "R-201"
- register: "R-356"
- evidence: "audits/REPORT-DRILL-backup-truth-2026-08-21.md"
verified:
date: 2026-08-09
date: 2026-08-22
verdict: contested
depth: source-read
+17
View File
@@ -1,3 +1,20 @@
## render_stands.py — the page stopped disagreeing with its own source (2026-08-22)
**The header's commit shas were hardcoded in the renderer, not read from the YAML.** `verified_on`
was parsed; `verified_against` never was. So the page printed
`felhom-agent 28ba8593b8, felhom-controller c732fe1283, hub 56f8aa611c` no matter what the dataset
said — and on 2026-08-22 the YAML header was updated to the current commits while the rendered page
went on citing the August 9th ones. **A build product silently disagreeing with the file it is built
from is exactly the failure this renderer's own docstring says it exists to prevent** ("it began going
stale the moment it was committed, which is the one thing a picture of 'where we stand' must not do").
`load()` now parses `verified_against` and the header renders from it; the literals are gone from both
the script and the output, checked.
**And the count beside it was wrong in a way that flattered us.** It read "N status(es) moved in that
pass" while N was every claim carrying a `changed:` block — 15, accumulated since the dataset began,
not 15 moves in one pass. It now reads "15 claim(s) carry a recorded status move (8 re-checked in this
pass)", the second number counting claims whose own `verified.date` equals the header date.
## due_checks_gate.py v1.0.1 + instructions_gate.py — the today-override announces itself (2026-08-18)
**Both gates read `FELHOM_GATE_TODAY` so their suites can control "today"; neither said so.** A shell
+17 -4
View File
@@ -45,6 +45,14 @@ def load(path):
line = raw.rstrip("\n")
if line.startswith("verified_on:"):
meta["date"] = line.split(":", 1)[1].strip()
# The commit shas the pass was verified against. Parsed rather than hardcoded: they were
# inlined in the header string until 2026-08-22, so the page kept printing the August 9th
# commits while the YAML said otherwise — a build product silently disagreeing with its own
# source, which is the exact failure this renderer exists to prevent.
if line.startswith(" ") and cur is None and ":" in line and not line.strip().startswith("-"):
k, _, v = line.strip().partition(":")
if k in ("felhom-agent", "felhom-controller", "hub") and v.strip():
meta.setdefault("against", {})[k] = v.strip()
if line.startswith(" - id:"):
cur = {"id": line.split(":", 1)[1].strip(), "sources": [], "changed": None}
claims.append(cur)
@@ -169,10 +177,15 @@ def main():
'<code>where-felhom-stands.yaml</code></b>, which cites the capability map, register row or '
'evidence document behind every claim — and which is checked by '
'<code>scripts/check_stands.py</code>.</div>')
o.append('<div style="color:#6b7a91;font-size:11.5px;margin-bottom:16px">Verified %s against '
'felhom-agent <code>28ba8593b8</code>, felhom-controller <code>c732fe1283</code>, hub '
'<code>56f8aa611c</code>. <b style="color:#fbbf24">%d status(es) moved in that pass</b> — '
'each is marked on the page.</div>' % (meta.get("date", "?"), len(changed)))
rechecked = sum(1 for c in claims if c.get("v_date") == meta.get("date"))
ag = meta.get("against", {})
against = ", ".join("%s <code>%s</code>" % (esc(k), esc(ag[k]))
for k in ("felhom-agent", "felhom-controller", "hub") if k in ag)
o.append('<div style="color:#6b7a91;font-size:11.5px;margin-bottom:16px">Verified %s%s. '
'<b style="color:#fbbf24">%d claim(s) carry a recorded status move</b> (%d re-checked '
'in this pass) — each is marked on the page.</div>'
% (esc(meta.get("date", "?")), (" against " + against) if against else "",
len(changed), rechecked))
leg = ['<div style="display:flex;gap:22px;flex-wrap:wrap;background:#121b2c;border:1px solid #1f2c44;'
'border-radius:6px;padding:11px 14px;margin-bottom:18px">']