diff --git a/documentation/architecture/where-felhom-stands.html b/documentation/architecture/where-felhom-stands.html index efdeb267..f5bc97fd 100644 --- a/documentation/architecture/where-felhom-stands.html +++ b/documentation/architecture/where-felhom-stands.html @@ -6,8 +6,8 @@

Where Felhom stands

What we built, what happens when things go wrong, and what is still missing. Generated from where-felhom-stands.yaml, which cites the capability map, register row or evidence document behind every claim — and which is checked by scripts/check_stands.py.
-
Verified 2026-08-09 against felhom-agent 28ba8593b8, felhom-controller c732fe1283, hub 56f8aa611c. 12 status(es) moved in that pass — each is marked on the page.
-
Walked (23)done end to end on real hardware, evidence on file
Built (14)shipped and tested, the real path never walked
Partial (14)some walked, some not — the note says which
Missing (4)does not exist
needs-hardwareonly a running box could settle it
+
Verified 2026-08-22 against felhom-agent 40d857b527, felhom-controller 2da259af38, hub 877fcd2a38. 15 claim(s) carry a recorded status move (8 re-checked in this pass) — each is marked on the page.
+
Walked (20)done end to end on real hardware, evidence on file
Built (14)shipped and tested, the real path never walked
Partial (17)some walked, some not — the note says which
Missing (4)does not exist
needs-hardwareonly a running box could settle it
The journey
Seven stages, left to right, as they happen to a person.
@@ -15,9 +15,9 @@
2  Making it theirs
An emailed one-time code; the customer sets their own password and the operator never sees itExercised live 2026-08-09: reset code accepted at /claim, 302 + session.capability-map Customer claim: one-time emailed code → customer sets own p… · evidence audits/REHEARSAL-byo-reinstall-2026-08-09.md
The customer can bind their own machine from a link, with no operator present — done once, for realattempts=0, locked=0, source customer_selfbind, 2026-07-18.capability-map Customer binds their own appliance (self-service) · evidence tests/VALIDATION-n100-rehearsal-2026-07-18.md
It has never been done by a person who is not the operatorMap records MISSING (as evidence). Still true after 2026-08-09.capability-map A customer (not the operator) performs a restore via UI alone
The same word is used for two different secrets across three surfaces; the email points at a page a rebuilt machine does not showWarning card. Cost a real code on 2026-08-09.register R-282 · register R-283
3  Using it
Apps installed from a catalog of about 52, with a memory guard and health-aware progress53 templates listed on the rebuilt box 2026-08-09; deploy driven live.capability-map Deploy an app from the catalog (env config, memory guard, h… · evidence audits/REHEARSAL-byo-reinstall-2026-08-09.md
Start, stop, restart, update, logs, remove — and the parts that must not be stopped cannot bechanged 2026-08-09, was walkedwhy it moved: no walk document cited by the map row or anywhere elsecapability-map App lifecycle: start/stop/restart/update/logs/remove/redeploy · register R-108
Reachable from anywhere through a tunnel, per-app addressesVerified tonight: the rebuilt box answered on its public URL from outside.capability-map Remote access via Cloudflare Tunnel + Traefik (per-app subd… · evidence audits/REHEARSAL-byo-reinstall-2026-08-09.md
Reachable on the home network when the internet is downneeds-hardwareOnly an observation on a running box with WAN pulled could settle it. Boxes are off.capability-map LAN access when internet is down (lan_resolver)
Phone photos, documents with text recognition, files from Windows Explorer or a Maccapability-map Files from Windows Explorer / Mac Finder (SMB server) · evidence audits/SPIKE-lan-discovery-2026-07-18.md
A one-tap launcher, and a read-only guest link for visitorscapability-map Indítópult (app launcher) — one-tap grid of the household's…
Media to a TVMap: MISSING.capability-map Media to TV via DLNA
Separate accounts per household memberMap: MISSING.capability-map Multiple household users / per-person accounts
4  Drives
A new drive is found, offered, formatted, mounted and enrolled — including on awkward older boot layoutsmoved twice: 2026-08-09 walked -> built, then 2026-08-10 back to walkedApplies to a NEW drive. Re-attaching an existing one after a reinstall is R-280 and fails.why it moved: receipt found 2026-08-10: a live throwaway drive on felhom-pve walked scan -> format -> mount -> PVE dir storage -> one-click Regisztralas, with the resulting storage.cfg entry and systemd mount unit recorded. The map already read PROVEN-LIVE; the dataset was behind it. Scope unchanged: this is a NEW drive. Re-attaching an existing one after a reinstall is R-280 and still fails.capability-map Drive wizard: scan/format/mount/enroll, incl. legacy-boot L… · register R-220 · evidence audits/SPIKE-raw-drive-enroll-2026-06-15.md
Moving data between drives, crash-safe; removing a drive safely; unplug detectedchanged 2026-08-09, was walkedwhy it moved: no walk document citedcapability-map Data migration between drives (all / per-app), crash-safe
A network drive can be browsed and hold bulk media, but may not hold an app's data — enforcedRefuseAsAppNamespace is the fail-closed predicate.capability-map Network storage (NAS) is browse + bulk-media only · register R-108 · evidence audits/R108-network-app-namespace-2026-07-30.md
After a reinstall the data drive cannot be re-attached through any dashboard routeWarning card. /api/disks/candidates returns empty; the restore page promises two clicks.register R-280 · evidence audits/REHEARSAL-byo-reinstall-2026-08-09.md
-
5  Backing up
App data on the machine, nightly database dumps, a copy on a second drivemoved twice: 2026-08-09 walked -> built, then 2026-08-10 back to walkedwhy it moved: receipt found 2026-08-10: an adversarial, destructive, unattended overnight campaign across both boxes and ep0, with A2 (one quiesce, two tiers) proven end to end. Map already read PROVEN-LIVE.capability-map Tier-2 secondary-drive copy: class-driven legs · evidence audits/CAMPAIGN-8-backup-restore-2026-07-27.md
A whole-machine archive that lands off the guest's own disk — a single-drive machine is recorded as degraded rather than pretendingchanged 2026-08-09, was walkedwhy it moved: no walk document citedcapability-map Whole-guest backup lands OFF the guest's own physical device · register R-165
An encrypted off-site copy, sealed with a key the operator cannot readVerified tonight from the repo itself: 18 snapshots, daily, unbroken.capability-map Offsite (restic → Hetzner Storage Box) · register R-199 · evidence audits/REHEARSAL-byo-reinstall-2026-08-09.md
The backups prove themselves: a restore is actually performed, unattended, on every tier, on both machineschanged 2026-08-09, was walked'on both machines, unattended, every tier' is a continuing claim about scheduled runs. Both boxes are off; the last recorded restore-test on demo-hp FAILED (notification_log 2026-08-05 restore_test_failed). Cannot be settled tonight.proof decay: PROOF-DECAY RULE FIRED (first time it has). A receipt EXISTS - architecture/_recovery-inventory-2026-07-28.md carries live journal lines for scheduled restore-tests on both boxes and both tiers - but it is superseded by later observation: demo-hp logged restore_test_failed on 2026-08-05, and the box has since been wiped and reinstalled (2026-08-09). The claim is about a CONTINUING scheduled behaviour, so a 2026-07-28 observation cannot carry it. Stays grey until a scheduled restore-test is seen passing on the rebuilt box. THE CAPABILITY MAP STILL READS PROVEN-LIVE (2026-08-03) AND IS NOW THE THING OUT OF STEP.why it moved: no walk document cited, and the last recorded restore-test on demo-hp FAILED (2026-08-05)capability-map Restore-proof is UNATTENDED — the scheduler covers EVERY tier · register R-86
The customer is warned before a drive fills, per drive, in their own languageR-177 (no operator-triggerable run) limits testing, not the capability.capability-map The customer is warned BEFORE a filesystem fills · register R-167 · evidence audits/SPIKE-r165-mp1-merge-2026-08-02.md
A backup that covered nothing still calls itself successfulWarning card.register R-240
+
5  Backing up
App data on the machine, nightly database dumps, a copy on a second drivemoved twice: 2026-08-09 walked -> built, then 2026-08-10 back to walkedTrue for every app but one, and that one was found on 2026-08-21: paperless-ngx's 72-table PostgreSQL was dumped nightly and written into a folder named after an app that does not exist, so it never entered the recovery unit, the second-drive copy or the off-site copy. Fixed and proven live 2026-08-22 (R-355) — the dump is now in the app's own unit and in the off-site snapshot for the first time. A catalogue-wide sweep, itself proven able to convict a planted case, says 1 of 53 was affected.why it moved: receipt found 2026-08-10: an adversarial, destructive, unattended overnight campaign across both boxes and ep0, with A2 (one quiesce, two tiers) proven end to end. Map already read PROVEN-LIVE.capability-map Tier-2 secondary-drive copy: class-driven legs · register R-355 · evidence audits/CAMPAIGN-8-backup-restore-2026-07-27.md · evidence audits/DRILL-backup-truth-2026-08-21/evidence/phase4-paperl…
A whole-machine archive that lands off the guest's own disk — a single-drive machine is recorded as degraded rather than pretendingchanged 2026-08-09, was walkedwhy it moved: no walk document citedcapability-map Whole-guest backup lands OFF the guest's own physical device · register R-165
An encrypted off-site copy, sealed with a key the operator cannot readchanged 2026-08-09, was walkedThe COPY is real and the key is still unreadable to us — what failed is DAILY and UNBROKEN. The 2026-08-09 note said '18 snapshots, daily, unbroken'; the next snapshot after 2026-08-09 08:30 was 2026-08-21 22:17, put there by hand during the drill. Twelve days, no alarm. Two causes in series: the 2026-08-21 rebuild lost the off-box target (R-193's shape), and after the self-heal restored it EVERY per-app switch was still off, so the first run logged 'backup OK: 0 app(s) backed up, 14s'.why it moved: the claim is about a CONTINUING daily copy; a 12-day silent gap was found on 2026-08-21 and nothing reported itcapability-map Offsite (restic → Hetzner Storage Box) · register R-199 · register R-193 · evidence audits/REHEARSAL-byo-reinstall-2026-08-09.md · evidence audits/REPORT-DRILL-backup-truth-2026-08-21.md
The backups prove themselves: a restore is actually performed, unattended, on every tier, on both machineschanged 2026-08-09, was walkedSTILL GREY, and for a sharper reason than in August. The scheduler runs and the check works — it failed again on 2026-08-21, unprompted, and named the cause exactly: after the reinstall the box presents a different PBS key than its own older archives were sealed with, so those archives cannot be opened at all (R-366). A tier whose archives are orphaned is a different alarm from a tier whose test failed, and only the second is being said.proof decay: PROOF-DECAY RULE FIRED (first time it has). A receipt EXISTS - architecture/_recovery-inventory-2026-07-28.md carries live journal lines for scheduled restore-tests on both boxes and both tiers - but it is superseded by later observation: demo-hp logged restore_test_failed on 2026-08-05, and the box has since been wiped and reinstalled (2026-08-09). The claim is about a CONTINUING scheduled behaviour, so a 2026-07-28 observation cannot carry it. Stays grey until a scheduled restore-test is seen passing on the rebuilt box. THE CAPABILITY MAP STILL READS PROVEN-LIVE (2026-08-03) AND IS NOW THE THING OUT OF STEP.why it moved: no walk document cited, and the last recorded restore-test on demo-hp FAILED (2026-08-05)capability-map Restore-proof is UNATTENDED — the scheduler covers EVERY tier · register R-86 · register R-366 · evidence audits/REPORT-DRILL-backup-truth-2026-08-21.md
The customer is warned before a drive fills, per drive, in their own languagechanged 2026-08-09, was walkedThe warning is real and was SEEN firing on 2026-08-21 with the right Hungarian copy, naming the drive and the free space. What 'BEFORE' cannot survive is the cadence: the watcher runs once a day at 03:30 plus once at startup (R-363), so a filesystem that fills at 03:31 goes unannounced for ~24 h. Watched live: the 69 GB volume carrying all 40-class app data was filled to 99% and the watcher said nothing, while the backup reserve was already refusing an app per run and telling the hub about it.why it moved: a daily check cannot carry the word BEFORE; observed silent for the whole window a filesystem sat at 99%capability-map The customer is warned BEFORE a filesystem fills · register R-167 · register R-363 · evidence audits/SPIKE-r165-mp1-merge-2026-08-02.md · evidence audits/REPORT-DRILL-backup-truth-2026-08-21.md
A backup that covered nothing still calls itself successfulWarning card — and the drill found two more of it, both live. (1) An off-site run with no app selected logs 'backup OK: 0 app(s) backed up'; the card does say 'nincs kijelölt alkalmazás' beside the green tick, so this one is honest if you read past the tick. (2) A restore that placed nothing reported success: '0 fájl visszaállítva', ok=true. The second is the one that matters and it is R-353, still open. The volume half of it is fixed (R-354): the message now names what came back.register R-240 · register R-353 · register R-354 · evidence audits/DRILL-backup-truth-2026-08-21/evidence/phase3-experi…
6  When something goes wrong
The machine watches itself and repairs some faults without telling anyone it had tochanged 2026-08-09, was walkedR-264 records that the self-heal counters reach the hub and are decoded nowhere.why it moved: no walk document citedcapability-map Box survives an unattended app or guest-network failure · register R-264
Failures reach the operator by email, one mail per run, every failing app namedchanged 2026-08-09, was walkedCONFIRMED against source: backup_run_failures digest is allowlisted (api/handler.go:1837), operator-only (dispatcher.go:423), templated (templates.go:48); recovery_unit_capture_failed is record-only (dispatcher.go:376). NOTE: R-182's register row still describes the PRE-FIX behaviour — see R-289.why it moved: the digest is wired and source-verified, but no run of it has been observed deliveringcapability-map A failed per-app Tier-1 backup reaches the OPERATOR — EVERY… · register R-182
Failures reach the customerThe customer leg is built; today's log shows customer-channel rows skipped as operator_only.capability-map App crashes → customer notified (one event per transition, …
An already-paired box is still told to pair itselfWarning card.register R-214 · register R-235
-
7  Getting everything back
A rebuilt machine shows a full-page recovery screen without anyone looking for it, and says plainly that nobody can replace a lost recovery codeSeen unsought on the rebuilt demo-hp 2026-08-09, seal date matching host_escrow.created_at.register R-193 · evidence audits/REHEARSAL-byo-reinstall-2026-08-09.md
The customer's code opens the sealed package and the data returns byte for byte — including accented Hungarian filenames, verified as raw bytesReproduced 2026-08-09: 4/4 byte-identical, name bytes NFC-preserved, out of snapshot 41c830db.register R-201 · evidence tests/walk5-r201-2026-08-07/journal.md · evidence audits/REHEARSAL-byo-reinstall-2026-08-09.md
Walked end to end with no command line inside the machine (2026-08-07)contestedCONTESTED-RESOLVED: not a contradiction. The walk proves the ROUTE needs no guest shell; the map's MISSING row is about a NON-OPERATOR doing it, which has still never happened. Two questions, one word 'customer'.register R-201 · evidence tests/walk5-r201-2026-08-07/journal.md · capability-map A customer (not the operator) performs a restore via UI alone
Putting restored files back where they belong is still manualWarning card. Confirmed 2026-08-09: the restore lands in a verification folder and says so.register R-213
The tripwire that says someone is opening this customer's backups does fireCONFIRMED tonight from the hub store: escrow_blob_served 2026-08-09 10:19:41Z = 12:19 CEST in the operator's mailbox. R-281's original claim of silence is WITHDRAWN.register R-281 · register R-285 · evidence audits/REHEARSAL-byo-reinstall-2026-08-09.md
+
7  Getting everything back
A rebuilt machine shows a full-page recovery screen without anyone looking for it, and says plainly that nobody can replace a lost recovery codeSeen unsought on the rebuilt demo-hp 2026-08-09, seal date matching host_escrow.created_at.register R-193 · evidence audits/REHEARSAL-byo-reinstall-2026-08-09.md
The customer's code opens the sealed package and the data returns byte for byte — including accented Hungarian filenames, verified as raw bytesStill true, and re-proven 2026-08-22 (5/5 byte-identical, both accented names as raw bytes) — but the SCOPE is narrower than the sentence sounds and was silently narrower still until v0.218.0. The drill found the off-site restore had NO named-volume leg at all: the archive sat in the unit, the snapshot and the checking folder and was never replayed, under a success message (R-354, fixed and proven 2026-08-22). AND 40 of the 53 catalogue apps STILL cannot run this route at all — it refuses first, saying a running app is not installed (R-356, open). So: proven for an app that declares a data drive; unproven and currently unreachable for the class whose entire dataset is a named volume.register R-201 · register R-354 · register R-356 · evidence tests/walk5-r201-2026-08-07/journal.md · evidence audits/REHEARSAL-byo-reinstall-2026-08-09.md · evidence audits/REPORT-DRILL-backup-truth-2026-08-21.md
Walked end to end with no command line inside the machine (2026-08-07)contestedCONTESTED-RESOLVED: not a contradiction. The walk proves the ROUTE needs no guest shell; the map's MISSING row is about a NON-OPERATOR doing it, which has still never happened. Two questions, one word 'customer'.register R-201 · evidence tests/walk5-r201-2026-08-07/journal.md · capability-map A customer (not the operator) performs a restore via UI alone
Putting restored files back where they belong is still manualWarning card. Confirmed 2026-08-09: the restore lands in a verification folder and says so.register R-213
The tripwire that says someone is opening this customer's backups does fireCONFIRMED tonight from the hub store: escrow_blob_served 2026-08-09 10:19:41Z = 12:19 CEST in the operator's mailbox. R-281's original claim of silence is WITHDRAWN.register R-281 · register R-285 · evidence audits/REHEARSAL-byo-reinstall-2026-08-09.md
When things go wrong
19 situations, and what actually happens in each.
@@ -28,7 +28,7 @@
An app crashes
An app crashes — the email leg has never been confirmed end to endConsistent with fault.operator-email: the digest is wired but its delivery is unobserved.capability-map App crashes → customer notified (one event per transition, …
Power cut mid-backup
Power cut mid-backupevidence audits/AUDIT-power-outage-recovery-2026-07-22.md
The guest is destroyed
The guest is destroyedregister R-201 · evidence audits/DRILL-r201-night-run-2026-08-04.md · evidence audits/DRILL-r201-night-run-2026-08-04.md
-
The whole machine is wiped and reinstalled
The whole machine is wiped and reinstalled — the data comes back4/4 byte-identical 2026-08-09.evidence audits/REHEARSAL-byo-reinstall-2026-08-09.md
+
The whole machine is wiped and reinstalled
The whole machine is wiped and reinstalled — the data comes backchanged 2026-08-09, was walkedThe 2026-08-09 rehearsal really did return 4/4 byte-identical, and that stands. What the next real reinstall showed (2026-08-21, demo-hp) is that the rebuild ORPHANS BOTH OFF-PREMISES TIERS at once, quietly: the restic target was lost and needed a self-heal plus a per-app re-enable before any copy resumed (R-193), and the PBS archives from before the reinstall cannot be opened by the rebuilt box at all, because it now presents a different key (R-366). The data came back in the rehearsal because the rehearsal restored it immediately; a machine left alone after a reinstall is not protected in the meantime and nothing says so.why it moved: a real reinstall on 2026-08-21 left both off-premises tiers broken - one silently for 12 days, the other unreadable - so 'the data comes back' holds only if someone restores it at onceregister R-193 · register R-366 · evidence audits/REHEARSAL-byo-reinstall-2026-08-09.md · evidence audits/REPORT-DRILL-backup-truth-2026-08-21.md
The whole machine is wiped and reinstalled
The whole machine is wiped and reinstalled — the journey needs a terminal twiceCONTESTED-RESOLVED against the map: the map's PROVEN-LIVE row is scoped to a controller-data-volume REBUILD (2026-08-04), not a whole-host reinstall. The map has no row for the host case, so there was no contradiction — only a gap.register R-273 · register R-280 · evidence audits/REHEARSAL-byo-reinstall-2026-08-09.md
The machine is stolen, and someone opens the backups
The machine is stolen, and someone opens the backups — the operator is toldConfirmed 2026-08-09 from the store and the mailbox.register R-281 · evidence audits/REHEARSAL-byo-reinstall-2026-08-09.md
The customer forgets their dashboard password
The customer forgets their dashboard passwordExercised 2026-08-09 via the reset code.evidence audits/REHEARSAL-byo-reinstall-2026-08-09.md
@@ -40,7 +40,7 @@
A fresh install picks up an old image
A fresh install picks up an old imageNarrowed 2026-08-09: the resume path fetched the vouched golden; a fresh install still takes the newest LOCAL archive with no manifest comparison.register R-274
The off-site provider account is deleted
The off-site provider account is deletedThe credential that holds the customer's documents can still delete.register R-95
The machine is switched off for an afternoon
The machine is switched off for an afternoon — no notion of expected downtimeConfirmed 2026-08-09: eight operator mails for deliberate, attended work.register R-285
-
A customer restores their own data with no help
A customer restores their own data with no helpcontestedCONTESTED-RESOLVED: the map's MISSING is about a NON-OPERATOR performing it; the walks prove the route, not the person. Neither record was wrong.capability-map A customer (not the operator) performs a restore via UI alone · register R-201
+
A customer restores their own data with no help
A customer restores their own data with no helpcontestedCONTESTED-RESOLVED: the map's MISSING is about a NON-OPERATOR performing it; the walks prove the route, not the person. Neither record was wrong. SINCE 2026-08-21 there is a second, harder blocker and it is not about the person at all: for the 40 of 53 apps that declare no data drive the off-site restore REFUSES before it starts, telling the customer a running app 'nincs telepítve' and to reinstall it to the same place — which those apps give them no way to choose (R-356). Those are exactly the apps whose whole dataset is a named volume. Until that is fixed, most customers cannot self-restore off-site even in principle.capability-map A customer (not the operator) performs a restore via UI alone · register R-201 · register R-356 · evidence audits/REPORT-DRILL-backup-truth-2026-08-21.md
Generated by scripts/render_stands.py from where-felhom-stands.yaml. Do not hand-edit this file — edit the data and regenerate. The 2026-08-09 React bundle is kept as where-felhom-stands-2026-08-09-snapshot.html and is not maintained.
\ No newline at end of file diff --git a/documentation/architecture/where-felhom-stands.yaml b/documentation/architecture/where-felhom-stands.yaml index 3f6b57ec..34588d5c 100644 --- a/documentation/architecture/where-felhom-stands.yaml +++ b/documentation/architecture/where-felhom-stands.yaml @@ -12,11 +12,11 @@ # verdict vocabulary: confirmed | downgraded | upgraded | contested | needs-hardware # depth: source-read (opened live source or evidence) | register+map (checked against the # register and capability map only) | needs-hardware (cannot be settled off-box) -verified_on: 2026-08-09 +verified_on: 2026-08-22 verified_against: - felhom-agent: 28ba8593b8 - felhom-controller: c732fe1283 - hub: 56f8aa611c + felhom-agent: 40d857b527 + felhom-controller: 2da259af38 + hub: 877fcd2a38 claims: - id: install.iso-selfregister band: journey @@ -301,17 +301,20 @@ claims: stage: 5 title: "App data on the machine, nightly database dumps, a copy on a second drive" status: walked + note: "True for every app but one, and that one was found on 2026-08-21: paperless-ngx's 72-table PostgreSQL was dumped nightly and written into a folder named after an app that does not exist, so it never entered the recovery unit, the second-drive copy or the off-site copy. Fixed and proven live 2026-08-22 (R-355) — the dump is now in the app's own unit and in the off-site snapshot for the first time. A catalogue-wide sweep, itself proven able to convict a planted case, says 1 of 53 was affected." sources: - capability-map: "Tier-2 secondary-drive copy: class-driven legs" + - register: "R-355" - evidence: "audits/CAMPAIGN-8-backup-restore-2026-07-27.md" + - evidence: "audits/DRILL-backup-truth-2026-08-21/evidence/phase4-paperless/FINDING.txt" changed: from: built also_moved: 2026-08-09 walked -> built reason_superseded: "no walk document cited" reason: "receipt found 2026-08-10: an adversarial, destructive, unattended overnight campaign across both boxes and ep0, with A2 (one quiesce, two tiers) proven end to end. Map already read PROVEN-LIVE." verified: - date: 2026-08-09 - verdict: upgraded + date: 2026-08-22 + verdict: confirmed depth: source-read - id: backup.whole-machine band: journey @@ -332,59 +335,75 @@ claims: band: journey stage: 5 title: "An encrypted off-site copy, sealed with a key the operator cannot read" - status: walked - note: "Verified tonight from the repo itself: 18 snapshots, daily, unbroken." + status: partial + note: "The COPY is real and the key is still unreadable to us — what failed is DAILY and UNBROKEN. The 2026-08-09 note said '18 snapshots, daily, unbroken'; the next snapshot after 2026-08-09 08:30 was 2026-08-21 22:17, put there by hand during the drill. Twelve days, no alarm. Two causes in series: the 2026-08-21 rebuild lost the off-box target (R-193's shape), and after the self-heal restored it EVERY per-app switch was still off, so the first run logged 'backup OK: 0 app(s) backed up, 14s'." sources: - capability-map: "Offsite (restic → Hetzner Storage Box)" - register: "R-199" + - register: "R-193" - evidence: "audits/REHEARSAL-byo-reinstall-2026-08-09.md" + - evidence: "audits/REPORT-DRILL-backup-truth-2026-08-21.md" + changed: + from: walked + reason: "the claim is about a CONTINUING daily copy; a 12-day silent gap was found on 2026-08-21 and nothing reported it" verified: - date: 2026-08-09 - verdict: confirmed + date: 2026-08-22 + verdict: downgraded depth: source-read - id: backup.restore-proof band: journey stage: 5 title: "The backups prove themselves: a restore is actually performed, unattended, on every tier, on both machines" status: built - note: "'on both machines, unattended, every tier' is a continuing claim about scheduled runs. Both boxes are off; the last recorded restore-test on demo-hp FAILED (notification_log 2026-08-05 restore_test_failed). Cannot be settled tonight." + note: "STILL GREY, and for a sharper reason than in August. The scheduler runs and the check works — it failed again on 2026-08-21, unprompted, and named the cause exactly: after the reinstall the box presents a different PBS key than its own older archives were sealed with, so those archives cannot be opened at all (R-366). A tier whose archives are orphaned is a different alarm from a tier whose test failed, and only the second is being said." sources: - capability-map: "Restore-proof is UNATTENDED — the scheduler covers EVERY tier" - register: "R-86" + - register: "R-366" + - evidence: "audits/REPORT-DRILL-backup-truth-2026-08-21.md" changed: from: walked reason: "no walk document cited, and the last recorded restore-test on demo-hp FAILED (2026-08-05)" decay: "PROOF-DECAY RULE FIRED (first time it has). A receipt EXISTS - architecture/_recovery-inventory-2026-07-28.md carries live journal lines for scheduled restore-tests on both boxes and both tiers - but it is superseded by later observation: demo-hp logged restore_test_failed on 2026-08-05, and the box has since been wiped and reinstalled (2026-08-09). The claim is about a CONTINUING scheduled behaviour, so a 2026-07-28 observation cannot carry it. Stays grey until a scheduled restore-test is seen passing on the rebuilt box. THE CAPABILITY MAP STILL READS PROVEN-LIVE (2026-08-03) AND IS NOW THE THING OUT OF STEP." + worse_2026_08_22: "It failed AGAIN, unprompted, on 2026-08-21 21:59 - and the cause is worse than 'untested'. Hub event 3016: the PBS archive of 2026-08-18 could not be restored because the manifest's key does not match the key the rebuilt box now presents. The archives that predate the 2026-08-21 reinstall are UNREADABLE to the machine that made them (R-366). Credit where due: the mechanism caught it and named the key mismatch precisely. The gap is that it is reported as 'a restore test failed' rather than 'your older whole-machine backups cannot be opened on this box'." verified: - date: 2026-08-09 + date: 2026-08-22 verdict: downgraded depth: needs-hardware - id: backup.fill-warning band: journey stage: 5 title: "The customer is warned before a drive fills, per drive, in their own language" - status: walked - note: "R-177 (no operator-triggerable run) limits testing, not the capability." + status: partial + note: "The warning is real and was SEEN firing on 2026-08-21 with the right Hungarian copy, naming the drive and the free space. What 'BEFORE' cannot survive is the cadence: the watcher runs once a day at 03:30 plus once at startup (R-363), so a filesystem that fills at 03:31 goes unannounced for ~24 h. Watched live: the 69 GB volume carrying all 40-class app data was filled to 99% and the watcher said nothing, while the backup reserve was already refusing an app per run and telling the hub about it." sources: - capability-map: "The customer is warned BEFORE a filesystem fills" - register: "R-167" + - register: "R-363" - evidence: "audits/SPIKE-r165-mp1-merge-2026-08-02.md" + - evidence: "audits/REPORT-DRILL-backup-truth-2026-08-21.md" + changed: + from: walked + reason: "a daily check cannot carry the word BEFORE; observed silent for the whole window a filesystem sat at 99%" verified: - date: 2026-08-09 - verdict: confirmed - depth: register+map + date: 2026-08-22 + verdict: downgraded + depth: source-read - id: backup.sikeres band: journey stage: 5 title: "A backup that covered nothing still calls itself successful" status: partial - note: "Warning card." + note: "Warning card — and the drill found two more of it, both live. (1) An off-site run with no app selected logs 'backup OK: 0 app(s) backed up'; the card does say 'nincs kijelölt alkalmazás' beside the green tick, so this one is honest if you read past the tick. (2) A restore that placed nothing reported success: '0 fájl visszaállítva', ok=true. The second is the one that matters and it is R-353, still open. The volume half of it is fixed (R-354): the message now names what came back." sources: - register: "R-240" + - register: "R-353" + - register: "R-354" + - evidence: "audits/DRILL-backup-truth-2026-08-21/evidence/phase3-experiment/messages-verbatim.txt" verified: - date: 2026-08-09 + date: 2026-08-22 verdict: confirmed - depth: register+map + depth: source-read - id: fault.selfheal band: journey stage: 6 @@ -460,13 +479,16 @@ claims: stage: 7 title: "The customer's code opens the sealed package and the data returns byte for byte — including accented Hungarian filenames, verified as raw bytes" status: walked - note: "Reproduced 2026-08-09: 4/4 byte-identical, name bytes NFC-preserved, out of snapshot 41c830db." + note: "Still true, and re-proven 2026-08-22 (5/5 byte-identical, both accented names as raw bytes) — but the SCOPE is narrower than the sentence sounds and was silently narrower still until v0.218.0. The drill found the off-site restore had NO named-volume leg at all: the archive sat in the unit, the snapshot and the checking folder and was never replayed, under a success message (R-354, fixed and proven 2026-08-22). AND 40 of the 53 catalogue apps STILL cannot run this route at all — it refuses first, saying a running app is not installed (R-356, open). So: proven for an app that declares a data drive; unproven and currently unreachable for the class whose entire dataset is a named volume." sources: - register: "R-201" + - register: "R-354" + - register: "R-356" - evidence: "tests/walk5-r201-2026-08-07/journal.md" - evidence: "audits/REHEARSAL-byo-reinstall-2026-08-09.md" + - evidence: "audits/REPORT-DRILL-backup-truth-2026-08-21.md" verified: - date: 2026-08-09 + date: 2026-08-22 verdict: confirmed depth: source-read - id: recover.no-shell @@ -581,13 +603,19 @@ claims: - id: fail.wiped-reinstalled.data band: failures title: "The whole machine is wiped and reinstalled — the data comes back" - status: walked - note: "4/4 byte-identical 2026-08-09." + status: partial + note: "The 2026-08-09 rehearsal really did return 4/4 byte-identical, and that stands. What the next real reinstall showed (2026-08-21, demo-hp) is that the rebuild ORPHANS BOTH OFF-PREMISES TIERS at once, quietly: the restic target was lost and needed a self-heal plus a per-app re-enable before any copy resumed (R-193), and the PBS archives from before the reinstall cannot be opened by the rebuilt box at all, because it now presents a different key (R-366). The data came back in the rehearsal because the rehearsal restored it immediately; a machine left alone after a reinstall is not protected in the meantime and nothing says so." sources: + - register: "R-193" + - register: "R-366" - evidence: "audits/REHEARSAL-byo-reinstall-2026-08-09.md" + - evidence: "audits/REPORT-DRILL-backup-truth-2026-08-21.md" + changed: + from: walked + reason: "a real reinstall on 2026-08-21 left both off-premises tiers broken - one silently for 12 days, the other unreadable - so 'the data comes back' holds only if someone restores it at once" verified: - date: 2026-08-09 - verdict: confirmed + date: 2026-08-22 + verdict: downgraded depth: source-read - id: fail.wiped-reinstalled.journey band: failures @@ -727,11 +755,13 @@ claims: band: failures title: "A customer restores their own data with no help" status: partial - note: "CONTESTED-RESOLVED: the map's MISSING is about a NON-OPERATOR performing it; the walks prove the route, not the person. Neither record was wrong." + note: "CONTESTED-RESOLVED: the map's MISSING is about a NON-OPERATOR performing it; the walks prove the route, not the person. Neither record was wrong. SINCE 2026-08-21 there is a second, harder blocker and it is not about the person at all: for the 40 of 53 apps that declare no data drive the off-site restore REFUSES before it starts, telling the customer a running app 'nincs telepítve' and to reinstall it to the same place — which those apps give them no way to choose (R-356). Those are exactly the apps whose whole dataset is a named volume. Until that is fixed, most customers cannot self-restore off-site even in principle." sources: - capability-map: "A customer (not the operator) performs a restore via UI alone" - register: "R-201" + - register: "R-356" + - evidence: "audits/REPORT-DRILL-backup-truth-2026-08-21.md" verified: - date: 2026-08-09 + date: 2026-08-22 verdict: contested depth: source-read diff --git a/scripts/CHANGELOG.md b/scripts/CHANGELOG.md index ae80fcc9..de76ed53 100644 --- a/scripts/CHANGELOG.md +++ b/scripts/CHANGELOG.md @@ -1,3 +1,20 @@ +## render_stands.py — the page stopped disagreeing with its own source (2026-08-22) + +**The header's commit shas were hardcoded in the renderer, not read from the YAML.** `verified_on` +was parsed; `verified_against` never was. So the page printed +`felhom-agent 28ba8593b8, felhom-controller c732fe1283, hub 56f8aa611c` no matter what the dataset +said — and on 2026-08-22 the YAML header was updated to the current commits while the rendered page +went on citing the August 9th ones. **A build product silently disagreeing with the file it is built +from is exactly the failure this renderer's own docstring says it exists to prevent** ("it began going +stale the moment it was committed, which is the one thing a picture of 'where we stand' must not do"). +`load()` now parses `verified_against` and the header renders from it; the literals are gone from both +the script and the output, checked. + +**And the count beside it was wrong in a way that flattered us.** It read "N status(es) moved in that +pass" while N was every claim carrying a `changed:` block — 15, accumulated since the dataset began, +not 15 moves in one pass. It now reads "15 claim(s) carry a recorded status move (8 re-checked in this +pass)", the second number counting claims whose own `verified.date` equals the header date. + ## due_checks_gate.py v1.0.1 + instructions_gate.py — the today-override announces itself (2026-08-18) **Both gates read `FELHOM_GATE_TODAY` so their suites can control "today"; neither said so.** A shell diff --git a/scripts/render_stands.py b/scripts/render_stands.py index 7ccc4df5..6fa1e9b2 100644 --- a/scripts/render_stands.py +++ b/scripts/render_stands.py @@ -45,6 +45,14 @@ def load(path): line = raw.rstrip("\n") if line.startswith("verified_on:"): meta["date"] = line.split(":", 1)[1].strip() + # The commit shas the pass was verified against. Parsed rather than hardcoded: they were + # inlined in the header string until 2026-08-22, so the page kept printing the August 9th + # commits while the YAML said otherwise — a build product silently disagreeing with its own + # source, which is the exact failure this renderer exists to prevent. + if line.startswith(" ") and cur is None and ":" in line and not line.strip().startswith("-"): + k, _, v = line.strip().partition(":") + if k in ("felhom-agent", "felhom-controller", "hub") and v.strip(): + meta.setdefault("against", {})[k] = v.strip() if line.startswith(" - id:"): cur = {"id": line.split(":", 1)[1].strip(), "sources": [], "changed": None} claims.append(cur) @@ -169,10 +177,15 @@ def main(): 'where-felhom-stands.yaml, which cites the capability map, register row or ' 'evidence document behind every claim — and which is checked by ' 'scripts/check_stands.py.') - o.append('
Verified %s against ' - 'felhom-agent 28ba8593b8, felhom-controller c732fe1283, hub ' - '56f8aa611c. %d status(es) moved in that pass — ' - 'each is marked on the page.
' % (meta.get("date", "?"), len(changed))) + rechecked = sum(1 for c in claims if c.get("v_date") == meta.get("date")) + ag = meta.get("against", {}) + against = ", ".join("%s %s" % (esc(k), esc(ag[k])) + for k in ("felhom-agent", "felhom-controller", "hub") if k in ag) + o.append('
Verified %s%s. ' + '%d claim(s) carry a recorded status move (%d re-checked ' + 'in this pass) — each is marked on the page.
' + % (esc(meta.get("date", "?")), (" against " + against) if against else "", + len(changed), rechecked)) leg = ['
']