Verify the standing picture against source: 12 downgrades, and the decay ran both ways
gates / gates (push) Successful in 21s
gates / gates (push) Successful in 21s
55 claims verified. Twelve moved, all downwards: walked 32 -> 20, built 5 -> 17. Register ceiling R-284 -> R-290. THE RULE DID NOT FIRE THE WAY IT WAS EXPECTED TO. Not one downgrade came from code moving under an old proof. All twelve came from step 1 of the same rule -- the cited evidence does not exist. Measured: of the 28 capability-map rows behind the page's claims, 8 carry a tests/ or audits/ path and 20 carry prose only. The green dots were drawn from rows that cite an argument, not a walk (R-290). The map, not the dataset, is what needs fixing -- it still says PROVEN-LIVE for all twelve. And once it ran backwards: fault.operator-email looked contradicted by R-182, but live source shows the backup_run_failures digest allowlisted, operator-only and templated, with recovery_unit_capture_failed now record-only. The claim is right and the REGISTER ROW is stale (R-289). The session went looking for stale proofs and found a stale defect. R-281 WITHDRAWN -- wrong in both directions, settled by the operator's mailbox. The tripwire DID fire (escrow_blob_served 10:19:41Z = 12:19 CEST) and false error-severity alarms fired too, for deliberate attended work (R-285). The measurement's cause is ESTABLISHED: the P7 query copied hub.db without hub.db-wal, and the signature is exact -- it reported "2 events all day, newest 00:30:07", and the rows at or before 00:30:07 number exactly 2. Timezone and wrong-key were tested and refuted. The control had been drawn from the same stale snapshot as the measurement, which is why it agreed (R-286). Part 4: NO WORKFLOW CHANGED, deliberately. The gate is not ref-sensitive -- it enumerates from the Gitea tags API, and both previous tag pushes passed. The red is TRUE: run 267 saw v0.120.0 downloadable, run 284 on the same commit saw 404. Who deleted the package is NOT established and is not guessed (R-287). The page is now generated from where-felhom-stands.yaml by scripts/render_stands.py: static, zero script tags, every moved status carrying a visible "changed, was X" chip. The React bundle -- whose content was gzip+base64 inside a JS module map -- is kept as a dated snapshot. scripts/check_stands.py gates the data and convicted 51 problems in my own first draft before the staged positive control ever ran.
This commit is contained in:
@@ -17,6 +17,12 @@ The Docker-only app-domain controller. Full per-area docs grounded in current so
|
||||
→ [`controller/README.md`](controller/README.md): module map, deploy & stack lifecycle, backup
|
||||
architecture, storage/monitoring/metrics, auth/hub/sync/integrations.
|
||||
|
||||
### Where we stand — `architecture/where-felhom-stands.*`
|
||||
The operator's one-page picture of what is proven, built, partial and missing.
|
||||
- [`architecture/where-felhom-stands.html`](architecture/where-felhom-stands.html) — **generated**; do not hand-edit
|
||||
- [`architecture/where-felhom-stands.yaml`](architecture/where-felhom-stands.yaml) — the data behind it; every claim cites its source. Gate: `scripts/check_stands.py`; regenerate with `scripts/render_stands.py`
|
||||
- [`architecture/where-felhom-stands-2026-08-09-snapshot.html`](architecture/where-felhom-stands-2026-08-09-snapshot.html) — **a dated snapshot, NOT maintained.** The original React bundle, kept for the record; its statuses are those of 2026-08-09 before the verification pass
|
||||
|
||||
### Host agent & platform — `architecture/`, `proxmox-platform.md`
|
||||
The operator-tier agent and the Proxmox platform.
|
||||
- [`architecture/01-topology-and-trust.md`](architecture/01-topology-and-trust.md) — topology & trust model
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -0,0 +1,727 @@
|
||||
# where-felhom-stands.yaml — the data behind documentation/architecture/where-felhom-stands.html
|
||||
#
|
||||
# THIS IS A VIEW, NEVER A SOURCE. Every entry cites the capability-map row, register row or
|
||||
# evidence document it derives from; an entry with no source is a defect, not a claim.
|
||||
# A status may not be RAISED here — if the evidence supports a stronger status than the
|
||||
# capability map records, the MAP changes first and this file follows it.
|
||||
# Regenerate the page after any status move: python3 scripts/render_stands.py
|
||||
#
|
||||
# YAML rather than JSON, deliberately: statuses move one line at a time and a YAML diff shows
|
||||
# which claim moved. A JSON re-dump reflows and shows the whole file.
|
||||
#
|
||||
# verdict vocabulary: confirmed | downgraded | upgraded | contested | needs-hardware
|
||||
# depth: source-read (opened live source or evidence) | register+map (checked against the
|
||||
# register and capability map only) | needs-hardware (cannot be settled off-box)
|
||||
verified_on: 2026-08-09
|
||||
verified_against:
|
||||
felhom-agent: 28ba8593b8
|
||||
felhom-controller: c732fe1283
|
||||
hub: 56f8aa611c
|
||||
claims:
|
||||
- id: install.iso-selfregister
|
||||
band: journey
|
||||
stage: 1
|
||||
title: "A blank machine installs itself from our own boot image and registers itself as unclaimed — proven on two different boards"
|
||||
status: walked
|
||||
note: "Two boards: N100 2026-07-18, HP t740 2026-07-21."
|
||||
sources:
|
||||
- capability-map: "Bare-metal Felhom ISO (blank hardware → zero-touch auto-install → first-boot host-install)"
|
||||
- evidence: "tests/VALIDATION-n100-rehearsal-2026-07-18.md"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: register+map
|
||||
- id: install.installer-by-tag
|
||||
band: journey
|
||||
stage: 1
|
||||
title: "The installer is published rather than pushed: rolling it back is one act"
|
||||
status: built
|
||||
note: "Gate 6 of hostinstall_gates asserts the manifest names an installer-v tag; ran green tonight."
|
||||
sources:
|
||||
- capability-map: "The installer is PUBLISHED, not pushed"
|
||||
- register: "R-110"
|
||||
changed:
|
||||
from: walked
|
||||
reason: "gate 6 asserts the manifest names an installer tag, but no walk of a rollback is on file"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: downgraded
|
||||
depth: source-read
|
||||
- id: install.byo
|
||||
band: journey
|
||||
stage: 1
|
||||
title: "Installing onto hardware the customer already owns — the path exists, the first real one has not happened"
|
||||
status: partial
|
||||
note: "UPGRADED: a real --mode byo install completed on demo-hp 2026-08-09 (Day-0 provision SUCCESS, 3m49s). Still not a customer's own hardware, so not 'walked' — but 'has not happened' is now false."
|
||||
sources:
|
||||
- evidence: "audits/REHEARSAL-byo-reinstall-2026-08-09.md"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: upgraded
|
||||
depth: source-read
|
||||
- id: install.nic-selfheal
|
||||
band: journey
|
||||
stage: 1
|
||||
title: "A machine that ends up on the wrong network port explains itself on screen and finds its way back — proven in a virtual drill, never on metal"
|
||||
status: partial
|
||||
note: "Map says the same: PROVEN-LIVE (nested drill — nested != metal)."
|
||||
sources:
|
||||
- capability-map: "Box survives a wrong-NIC install"
|
||||
- evidence: "audits/SPIKE-firstboot-nic-sweep-2026-07-22.md"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: register+map
|
||||
- id: install.reinstall-refuses
|
||||
band: journey
|
||||
stage: 1
|
||||
title: "Reinstalling a machine we previously installed refuses, because our own removal leaves a name service holding the port our own installer checks"
|
||||
status: partial
|
||||
note: "Warning card. Confirmed live 2026-08-09; dnsmasq restarted by --uninstall seizes :53."
|
||||
sources:
|
||||
- register: "R-272"
|
||||
- evidence: "audits/REHEARSAL-byo-reinstall-2026-08-09.md"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: source-read
|
||||
- id: claim.one-time-code
|
||||
band: journey
|
||||
stage: 2
|
||||
title: "An emailed one-time code; the customer sets their own password and the operator never sees it"
|
||||
status: walked
|
||||
note: "Exercised live 2026-08-09: reset code accepted at /claim, 302 + session."
|
||||
sources:
|
||||
- capability-map: "Customer claim: one-time emailed code → customer sets own password"
|
||||
- evidence: "audits/REHEARSAL-byo-reinstall-2026-08-09.md"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: source-read
|
||||
- id: claim.selfbind
|
||||
band: journey
|
||||
stage: 2
|
||||
title: "The customer can bind their own machine from a link, with no operator present — done once, for real"
|
||||
status: walked
|
||||
note: "attempts=0, locked=0, source customer_selfbind, 2026-07-18."
|
||||
sources:
|
||||
- capability-map: "Customer binds their own appliance (self-service)"
|
||||
- evidence: "tests/VALIDATION-n100-rehearsal-2026-07-18.md"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: register+map
|
||||
- id: claim.never-by-non-operator
|
||||
band: journey
|
||||
stage: 2
|
||||
title: "It has never been done by a person who is not the operator"
|
||||
status: partial
|
||||
note: "Map records MISSING (as evidence). Still true after 2026-08-09."
|
||||
sources:
|
||||
- capability-map: "A customer (not the operator) performs a restore via UI alone"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: register+map
|
||||
- id: claim.code-naming
|
||||
band: journey
|
||||
stage: 2
|
||||
title: "The same word is used for two different secrets across three surfaces; the email points at a page a rebuilt machine does not show"
|
||||
status: partial
|
||||
note: "Warning card. Cost a real code on 2026-08-09."
|
||||
sources:
|
||||
- register: "R-282"
|
||||
- register: "R-283"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: source-read
|
||||
- id: use.catalog
|
||||
band: journey
|
||||
stage: 3
|
||||
title: "Apps installed from a catalog of about 52, with a memory guard and health-aware progress"
|
||||
status: walked
|
||||
note: "53 templates listed on the rebuilt box 2026-08-09; deploy driven live."
|
||||
sources:
|
||||
- capability-map: "Deploy an app from the catalog (env config, memory guard, health-aware progress)"
|
||||
- evidence: "audits/REHEARSAL-byo-reinstall-2026-08-09.md"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: source-read
|
||||
- id: use.lifecycle
|
||||
band: journey
|
||||
stage: 3
|
||||
title: "Start, stop, restart, update, logs, remove — and the parts that must not be stopped cannot be"
|
||||
status: built
|
||||
sources:
|
||||
- capability-map: "App lifecycle: start/stop/restart/update/logs/remove/redeploy"
|
||||
- register: "R-108"
|
||||
changed:
|
||||
from: walked
|
||||
reason: "no walk document cited by the map row or anywhere else"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: downgraded
|
||||
depth: register+map
|
||||
- id: use.tunnel
|
||||
band: journey
|
||||
stage: 3
|
||||
title: "Reachable from anywhere through a tunnel, per-app addresses"
|
||||
status: walked
|
||||
note: "Verified tonight: the rebuilt box answered on its public URL from outside."
|
||||
sources:
|
||||
- capability-map: "Remote access via Cloudflare Tunnel + Traefik (per-app subdomains)"
|
||||
- evidence: "audits/REHEARSAL-byo-reinstall-2026-08-09.md"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: source-read
|
||||
- id: use.lan-fallback
|
||||
band: journey
|
||||
stage: 3
|
||||
title: "Reachable on the home network when the internet is down"
|
||||
status: built
|
||||
note: "Only an observation on a running box with WAN pulled could settle it. Boxes are off."
|
||||
sources:
|
||||
- capability-map: "LAN access when internet is down (lan_resolver)"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: needs-hardware
|
||||
depth: needs-hardware
|
||||
- id: use.files
|
||||
band: journey
|
||||
stage: 3
|
||||
title: "Phone photos, documents with text recognition, files from Windows Explorer or a Mac"
|
||||
status: walked
|
||||
sources:
|
||||
- capability-map: "Files from Windows Explorer / Mac Finder (SMB server)"
|
||||
- evidence: "audits/SPIKE-lan-discovery-2026-07-18.md"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: register+map
|
||||
- id: use.launcher
|
||||
band: journey
|
||||
stage: 3
|
||||
title: "A one-tap launcher, and a read-only guest link for visitors"
|
||||
status: built
|
||||
sources:
|
||||
- capability-map: "Indítópult (app launcher) — one-tap grid of the household's openable apps"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: register+map
|
||||
- id: use.dlna
|
||||
band: journey
|
||||
stage: 3
|
||||
title: "Media to a TV"
|
||||
status: missing
|
||||
note: "Map: MISSING."
|
||||
sources:
|
||||
- capability-map: "Media to TV via DLNA"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: register+map
|
||||
- id: use.multiuser
|
||||
band: journey
|
||||
stage: 3
|
||||
title: "Separate accounts per household member"
|
||||
status: missing
|
||||
note: "Map: MISSING."
|
||||
sources:
|
||||
- capability-map: "Multiple household users / per-person accounts"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: register+map
|
||||
- id: drives.enrol
|
||||
band: journey
|
||||
stage: 4
|
||||
title: "A new drive is found, offered, formatted, mounted and enrolled — including on awkward older boot layouts"
|
||||
status: built
|
||||
note: "Applies to a NEW drive. Re-attaching an existing one after a reinstall is R-280 and fails."
|
||||
sources:
|
||||
- capability-map: "Drive wizard: scan/format/mount/enroll, incl. legacy-boot LVM-root hosts"
|
||||
- register: "R-220"
|
||||
changed:
|
||||
from: walked
|
||||
reason: "the 2026-08-09 walk exercised RE-attach (which failed, R-280); first-enrolment of a NEW drive has no walk on file"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: downgraded
|
||||
depth: source-read
|
||||
- id: drives.migrate
|
||||
band: journey
|
||||
stage: 4
|
||||
title: "Moving data between drives, crash-safe; removing a drive safely; unplug detected"
|
||||
status: built
|
||||
sources:
|
||||
- capability-map: "Data migration between drives (all / per-app), crash-safe"
|
||||
changed:
|
||||
from: walked
|
||||
reason: "no walk document cited"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: downgraded
|
||||
depth: register+map
|
||||
- id: drives.nas
|
||||
band: journey
|
||||
stage: 4
|
||||
title: "A network drive can be browsed and hold bulk media, but may not hold an app's data — enforced"
|
||||
status: walked
|
||||
note: "RefuseAsAppNamespace is the fail-closed predicate."
|
||||
sources:
|
||||
- capability-map: "Network storage (NAS) is browse + bulk-media only"
|
||||
- register: "R-108"
|
||||
- evidence: "audits/R108-network-app-namespace-2026-07-30.md"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: register+map
|
||||
- id: drives.reattach-wall
|
||||
band: journey
|
||||
stage: 4
|
||||
title: "After a reinstall the data drive cannot be re-attached through any dashboard route"
|
||||
status: partial
|
||||
note: "Warning card. /api/disks/candidates returns empty; the restore page promises two clicks."
|
||||
sources:
|
||||
- register: "R-280"
|
||||
- evidence: "audits/REHEARSAL-byo-reinstall-2026-08-09.md"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: source-read
|
||||
- id: backup.tier1
|
||||
band: journey
|
||||
stage: 5
|
||||
title: "App data on the machine, nightly database dumps, a copy on a second drive"
|
||||
status: built
|
||||
sources:
|
||||
- capability-map: "Tier-2 secondary-drive copy: class-driven legs"
|
||||
changed:
|
||||
from: walked
|
||||
reason: "no walk document cited"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: downgraded
|
||||
depth: register+map
|
||||
- id: backup.whole-machine
|
||||
band: journey
|
||||
stage: 5
|
||||
title: "A whole-machine archive that lands off the guest's own disk — a single-drive machine is recorded as degraded rather than pretending"
|
||||
status: built
|
||||
sources:
|
||||
- capability-map: "Whole-guest backup lands OFF the guest's own physical device"
|
||||
- register: "R-165"
|
||||
changed:
|
||||
from: walked
|
||||
reason: "no walk document cited"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: downgraded
|
||||
depth: register+map
|
||||
- id: backup.offsite
|
||||
band: journey
|
||||
stage: 5
|
||||
title: "An encrypted off-site copy, sealed with a key the operator cannot read"
|
||||
status: walked
|
||||
note: "Verified tonight from the repo itself: 18 snapshots, daily, unbroken."
|
||||
sources:
|
||||
- capability-map: "Offsite (restic → Hetzner Storage Box)"
|
||||
- register: "R-199"
|
||||
- evidence: "audits/REHEARSAL-byo-reinstall-2026-08-09.md"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: source-read
|
||||
- id: backup.restore-proof
|
||||
band: journey
|
||||
stage: 5
|
||||
title: "The backups prove themselves: a restore is actually performed, unattended, on every tier, on both machines"
|
||||
status: built
|
||||
note: "'on both machines, unattended, every tier' is a continuing claim about scheduled runs. Both boxes are off; the last recorded restore-test on demo-hp FAILED (notification_log 2026-08-05 restore_test_failed). Cannot be settled tonight."
|
||||
sources:
|
||||
- capability-map: "Restore-proof is UNATTENDED — the scheduler covers EVERY tier"
|
||||
- register: "R-86"
|
||||
changed:
|
||||
from: walked
|
||||
reason: "no walk document cited, and the last recorded restore-test on demo-hp FAILED (2026-08-05)"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: downgraded
|
||||
depth: needs-hardware
|
||||
- id: backup.fill-warning
|
||||
band: journey
|
||||
stage: 5
|
||||
title: "The customer is warned before a drive fills, per drive, in their own language"
|
||||
status: walked
|
||||
note: "R-177 (no operator-triggerable run) limits testing, not the capability."
|
||||
sources:
|
||||
- capability-map: "The customer is warned BEFORE a filesystem fills"
|
||||
- register: "R-167"
|
||||
- evidence: "audits/SPIKE-r165-mp1-merge-2026-08-02.md"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: register+map
|
||||
- id: backup.sikeres
|
||||
band: journey
|
||||
stage: 5
|
||||
title: "A backup that covered nothing still calls itself successful"
|
||||
status: partial
|
||||
note: "Warning card."
|
||||
sources:
|
||||
- register: "R-240"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: register+map
|
||||
- id: fault.selfheal
|
||||
band: journey
|
||||
stage: 6
|
||||
title: "The machine watches itself and repairs some faults without telling anyone it had to"
|
||||
status: built
|
||||
note: "R-264 records that the self-heal counters reach the hub and are decoded nowhere."
|
||||
sources:
|
||||
- capability-map: "Box survives an unattended app or guest-network failure"
|
||||
- register: "R-264"
|
||||
changed:
|
||||
from: walked
|
||||
reason: "no walk document cited"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: downgraded
|
||||
depth: source-read
|
||||
- id: fault.operator-email
|
||||
band: journey
|
||||
stage: 6
|
||||
title: "Failures reach the operator by email, one mail per run, every failing app named"
|
||||
status: built
|
||||
note: "CONFIRMED against source: backup_run_failures digest is allowlisted (api/handler.go:1837), operator-only (dispatcher.go:423), templated (templates.go:48); recovery_unit_capture_failed is record-only (dispatcher.go:376). NOTE: R-182's register row still describes the PRE-FIX behaviour — see R-289."
|
||||
sources:
|
||||
- capability-map: "A failed per-app Tier-1 backup reaches the OPERATOR — EVERY failing app, in ONE mail per run"
|
||||
- register: "R-182"
|
||||
changed:
|
||||
from: walked
|
||||
reason: "the digest is wired and source-verified, but no run of it has been observed delivering"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: downgraded
|
||||
depth: source-read
|
||||
- id: fault.customer-email
|
||||
band: journey
|
||||
stage: 6
|
||||
title: "Failures reach the customer"
|
||||
status: built
|
||||
note: "The customer leg is built; today's log shows customer-channel rows skipped as operator_only."
|
||||
sources:
|
||||
- capability-map: "App crashes → customer notified (one event per transition, no flapping spam)"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: source-read
|
||||
- id: fault.already-paired
|
||||
band: journey
|
||||
stage: 6
|
||||
title: "An already-paired box is still told to pair itself"
|
||||
status: partial
|
||||
note: "Warning card."
|
||||
sources:
|
||||
- register: "R-214"
|
||||
- register: "R-235"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: register+map
|
||||
- id: recover.screen
|
||||
band: journey
|
||||
stage: 7
|
||||
title: "A rebuilt machine shows a full-page recovery screen without anyone looking for it, and says plainly that nobody can replace a lost recovery code"
|
||||
status: walked
|
||||
note: "Seen unsought on the rebuilt demo-hp 2026-08-09, seal date matching host_escrow.created_at."
|
||||
sources:
|
||||
- register: "R-193"
|
||||
- evidence: "audits/REHEARSAL-byo-reinstall-2026-08-09.md"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: source-read
|
||||
- id: recover.byte-identical
|
||||
band: journey
|
||||
stage: 7
|
||||
title: "The customer's code opens the sealed package and the data returns byte for byte — including accented Hungarian filenames, verified as raw bytes"
|
||||
status: walked
|
||||
note: "Reproduced 2026-08-09: 4/4 byte-identical, name bytes NFC-preserved, out of snapshot 41c830db."
|
||||
sources:
|
||||
- register: "R-201"
|
||||
- evidence: "tests/walk5-r201-2026-08-07/journal.md"
|
||||
- evidence: "audits/REHEARSAL-byo-reinstall-2026-08-09.md"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: source-read
|
||||
- id: recover.no-shell
|
||||
band: journey
|
||||
stage: 7
|
||||
title: "Walked end to end with no command line inside the machine (2026-08-07)"
|
||||
status: walked
|
||||
note: "CONTESTED-RESOLVED: not a contradiction. The walk proves the ROUTE needs no guest shell; the map's MISSING row is about a NON-OPERATOR doing it, which has still never happened. Two questions, one word 'customer'."
|
||||
sources:
|
||||
- register: "R-201"
|
||||
- evidence: "tests/walk5-r201-2026-08-07/journal.md"
|
||||
- capability-map: "A customer (not the operator) performs a restore via UI alone"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: contested
|
||||
depth: source-read
|
||||
- id: recover.putback
|
||||
band: journey
|
||||
stage: 7
|
||||
title: "Putting restored files back where they belong is still manual"
|
||||
status: partial
|
||||
note: "Warning card. Confirmed 2026-08-09: the restore lands in a verification folder and says so."
|
||||
sources:
|
||||
- register: "R-213"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: source-read
|
||||
- id: recover.tripwire
|
||||
band: journey
|
||||
stage: 7
|
||||
title: "The tripwire that says someone is opening this customer's backups does fire"
|
||||
status: walked
|
||||
note: "CONFIRMED tonight from the hub store: escrow_blob_served 2026-08-09 10:19:41Z = 12:19 CEST in the operator's mailbox. R-281's original claim of silence is WITHDRAWN."
|
||||
sources:
|
||||
- register: "R-281"
|
||||
- register: "R-285"
|
||||
- evidence: "audits/REHEARSAL-byo-reinstall-2026-08-09.md"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: source-read
|
||||
- id: fail.disk-failing
|
||||
band: failures
|
||||
title: "A disk starts failing — healthy path only; a genuinely failing disk has never been seen"
|
||||
status: partial
|
||||
note: "Only a failing disk on a running box could settle it."
|
||||
sources:
|
||||
- capability-map: "Lemez-egészség felügyelet: per-disk SMART kártya"
|
||||
- evidence: "audits/SPIKE-smart-coverage-2026-07-25.md"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: needs-hardware
|
||||
depth: needs-hardware
|
||||
- id: fail.backup-drive-unplugged
|
||||
band: failures
|
||||
title: "The backup drive is unplugged"
|
||||
status: walked
|
||||
sources:
|
||||
- capability-map: "An ABSENT backup-target drive raises its OWN alarm"
|
||||
- evidence: "audits/R116-v0116-2026-07-30.md"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: register+map
|
||||
- id: fail.drive-filling
|
||||
band: failures
|
||||
title: "A drive is filling up"
|
||||
status: built
|
||||
sources:
|
||||
- register: "R-167"
|
||||
changed:
|
||||
from: walked
|
||||
reason: "no walk document cited"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: downgraded
|
||||
depth: register+map
|
||||
- id: fail.app-crash
|
||||
band: failures
|
||||
title: "An app crashes — the email leg has never been confirmed end to end"
|
||||
status: built
|
||||
note: "Consistent with fault.operator-email: the digest is wired but its delivery is unobserved."
|
||||
sources:
|
||||
- capability-map: "App crashes → customer notified (one event per transition, no flapping spam)"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: source-read
|
||||
- id: fail.power-cut
|
||||
band: failures
|
||||
title: "Power cut mid-backup"
|
||||
status: walked
|
||||
sources:
|
||||
- evidence: "audits/AUDIT-power-outage-recovery-2026-07-22.md"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: register+map
|
||||
- id: fail.guest-destroyed
|
||||
band: failures
|
||||
title: "The guest is destroyed"
|
||||
status: walked
|
||||
sources:
|
||||
- register: "R-201"
|
||||
- evidence: "audits/DRILL-r201-night-run-2026-08-04.md"
|
||||
- evidence: "audits/DRILL-r201-night-run-2026-08-04.md"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: register+map
|
||||
- id: fail.wiped-reinstalled.data
|
||||
band: failures
|
||||
title: "The whole machine is wiped and reinstalled — the data comes back"
|
||||
status: walked
|
||||
note: "4/4 byte-identical 2026-08-09."
|
||||
sources:
|
||||
- evidence: "audits/REHEARSAL-byo-reinstall-2026-08-09.md"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: source-read
|
||||
- id: fail.wiped-reinstalled.journey
|
||||
band: failures
|
||||
title: "The whole machine is wiped and reinstalled — the journey needs a terminal twice"
|
||||
status: partial
|
||||
note: "CONTESTED-RESOLVED against the map: the map's PROVEN-LIVE row is scoped to a controller-data-volume REBUILD (2026-08-04), not a whole-host reinstall. The map has no row for the host case, so there was no contradiction — only a gap."
|
||||
sources:
|
||||
- register: "R-273"
|
||||
- register: "R-280"
|
||||
- evidence: "audits/REHEARSAL-byo-reinstall-2026-08-09.md"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: source-read
|
||||
- id: fail.stolen-machine
|
||||
band: failures
|
||||
title: "The machine is stolen, and someone opens the backups — the operator is told"
|
||||
status: walked
|
||||
note: "Confirmed 2026-08-09 from the store and the mailbox."
|
||||
sources:
|
||||
- register: "R-281"
|
||||
- evidence: "audits/REHEARSAL-byo-reinstall-2026-08-09.md"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: source-read
|
||||
- id: fail.forgot-password
|
||||
band: failures
|
||||
title: "The customer forgets their dashboard password"
|
||||
status: walked
|
||||
note: "Exercised 2026-08-09 via the reset code."
|
||||
sources:
|
||||
- evidence: "audits/REHEARSAL-byo-reinstall-2026-08-09.md"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: source-read
|
||||
- id: fail.lost-recovery-code
|
||||
band: failures
|
||||
title: "The customer loses their recovery code — by design, the data is unrecoverable"
|
||||
status: built
|
||||
note: "The recovery screen states it in Hungarian."
|
||||
sources:
|
||||
- capability-map: "Escrow ceremony: customer-facing wizard, one-shot R claim, operator zero-knowledge"
|
||||
- register: "R-198"
|
||||
changed:
|
||||
from: walked
|
||||
reason: "a by-design refusal; no walk document cited"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: downgraded
|
||||
depth: source-read
|
||||
- id: fail.moves-house
|
||||
band: failures
|
||||
title: "The machine moves house / new network"
|
||||
status: walked
|
||||
sources:
|
||||
- capability-map: "Box survives a site/network change (relocation, different subnet, DHCP re-lease)"
|
||||
- evidence: "audits/AUDIT-vacation-remote-ops-2026-07-20.md"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: register+map
|
||||
- id: fail.internet-down
|
||||
band: failures
|
||||
title: "The internet goes down — built, never walked"
|
||||
status: built
|
||||
note: "Needs a running box with WAN pulled."
|
||||
sources:
|
||||
- capability-map: "LAN access when internet is down (lan_resolver)"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: needs-hardware
|
||||
depth: needs-hardware
|
||||
- id: fail.hub-down
|
||||
band: failures
|
||||
title: "The hub is down"
|
||||
status: built
|
||||
sources:
|
||||
- capability-map: "Config/state change round-trips in seconds (hub↔box immediacy"
|
||||
changed:
|
||||
from: walked
|
||||
reason: "no walk document cited"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: downgraded
|
||||
depth: register+map
|
||||
- id: fail.broken-release
|
||||
band: failures
|
||||
title: "We ship a broken release — the guard is missing"
|
||||
status: missing
|
||||
note: "Proven the hard way on 2026-08-09: a vouched agent version had no git tag and every install died at 5/8. Both guards still owed."
|
||||
sources:
|
||||
- register: "R-273"
|
||||
- register: "R-287"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: source-read
|
||||
- id: fail.stale-image
|
||||
band: failures
|
||||
title: "A fresh install picks up an old image"
|
||||
status: partial
|
||||
note: "Narrowed 2026-08-09: the resume path fetched the vouched golden; a fresh install still takes the newest LOCAL archive with no manifest comparison."
|
||||
sources:
|
||||
- register: "R-274"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: source-read
|
||||
- id: fail.offsite-account-deleted
|
||||
band: failures
|
||||
title: "The off-site provider account is deleted"
|
||||
status: partial
|
||||
note: "The credential that holds the customer's documents can still delete."
|
||||
sources:
|
||||
- register: "R-95"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: register+map
|
||||
- id: fail.expected-downtime
|
||||
band: failures
|
||||
title: "The machine is switched off for an afternoon — no notion of expected downtime"
|
||||
status: missing
|
||||
note: "Confirmed 2026-08-09: eight operator mails for deliberate, attended work."
|
||||
sources:
|
||||
- register: "R-285"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: confirmed
|
||||
depth: source-read
|
||||
- id: fail.customer-self-restore
|
||||
band: failures
|
||||
title: "A customer restores their own data with no help"
|
||||
status: partial
|
||||
note: "CONTESTED-RESOLVED: the map's MISSING is about a NON-OPERATOR performing it; the walks prove the route, not the person. Neither record was wrong."
|
||||
sources:
|
||||
- capability-map: "A customer (not the operator) performs a restore via UI alone"
|
||||
- register: "R-201"
|
||||
verified:
|
||||
date: 2026-08-09
|
||||
verdict: contested
|
||||
depth: source-read
|
||||
@@ -509,10 +509,16 @@ applied.** The one that matters: Scenario A **fails against today's tree** with
|
||||
| **R-278** | **demo-felhom's off-site tier has never completed a run and has been stuck for six days.** `offsite.state=needs_credential` since the 2026-08-03 guest rebuild; the hub's own alarm reads *"enabled + escrowed but no run has EVER succeeded"*; the controller's `offsite-credential-retry` job runs every 5 minutes and completes in 0 s, doing nothing. R-193's fix (the recovery SCREEN, controller 0.200.0) is present on the box, so the remedy exists — it just needs the customer-present ceremony that nobody has run, which is R-243's shape (*"a machine waiting for its recovery code can stop backing up off-site without alarming us"*) landing on a real box. **Contrast that makes it a defect and not a chore:** demo-hp, same rebuild, same day, recovered and has 18 snapshots | **READY (S) — NEW 2026-08-09** | — | Either the self-heal reconciler owns this shape end-to-end, or the box must say plainly on the dashboard that it is unprotected pending the recovery code | CC |
|
||||
| **R-279** | **There is no operator-triggerable off-site backup.** The only route to `POST /backup/offbox/run` is the customer's own dashboard session; `signed_jobs` carries opaque operator-SIGNED blobs and the hub holds no signing key. This cost the rehearsal a stop: preparing the run needed one off-site push and there was no operator path to it. Sibling of **R-177** (no operator-triggerable fill check) | **READY (XS) — NEW 2026-08-09** | — | Same shape as R-177; solve both together | CC |
|
||||
| **R-280** | **RANK 1 — after a reinstall the data drive cannot be re-attached through ANY dashboard route, and the restore page promises it is "two clicks".** Measured on the rebuilt demo-hp, 2026-08-09. The restore page diagnoses the situation perfectly and then sends the customer to an empty page: *„Előbb csatold vissza az adatmeghajtót. A mentéseid megvannak, és a meghajtók is megvannak — újratelepítés után viszont a gép még nem ismeri őket, ezért most nincs hová visszaállítani. **Ez két kattintás:** Tárhely → Meghajtók, »Meglévő meghajtó csatolása«."* **It is not two clicks; it is zero possible clicks.** `GET /api/disks/candidates` → `{"initialize":[],"attach":[]}`, so both wizards render an empty selector, and `Tárhely → Meghajtók` reads „Nincs regisztrált adattároló" with an empty unregistered list. **The agent is not at fault** — `GET /api/disks` returns the NVMe in full (1.0 TB, SMART PASSED, `mount_path:/mnt/nvme-1tb`, `guest_attached:false`), so the channel and enumeration work. **ROOT CAUSE:** `handleDiskCandidates` builds both lists from `ListCandidateDisks`, the UNCLAIMED-disk scan; demo-hp's NVMe is deliberately BOTH the user-data drive and the `felhom-backup` target (`operations/nodes.md`), so it is claimed and never offered. That filter is **correct for `initialize`** (never offer to format a disk in use — `/storage/init` even says so: *„Rendszer- és biztonsági-mentés meghajtók itt nem jelennek meg — azok védettek"*) and **over-broad for `attach`**, which is non-destructive by definition and whose own page says *„A meghajtón lévő adatok nem törlődnek — a csatolás csak elérhetővé teszi azokat."* **It cascades:** no store → Calibre-Web's install page degrades to *„Nincs regisztrált adattároló — adja meg kézzel az útvonalat"* and demands a hand-typed `E-könyvtár útvonal`; no app → the restore rows read „Nincs telepítve". **THE ESCAPE HATCH WORKS AND NO CUSTOMER COULD FIND IT:** `POST /settings/storage/add` with `storage_path=/mnt/sys_drive` succeeded first try (*„Adattároló sikeresen hozzáadva"*) — and `/mnt/sys_drive` is an internal path, the very one registered before the wipe. Once registered, everything unblocked and the deploy form became a proper picker (*„Tárhely (sys_drive) — 64.2 GB szabad"*). **This is R-220's successor:** R-220 was closed as "drives unenrollable after a rebuild — fixed"; enumeration is fixed, OFFERING is not | **READY (M) — NEW 2026-08-09** | — | Populate `attach` from mounted-but-unregistered filesystems rather than from the unclaimed-DISK scan; and never print "two clicks" without asserting the destination is non-empty | CC |
|
||||
| **R-281** | **The hub said NOTHING through an entire reinstall — and the tripwire for a sealed-backup unseal did not fire on a real unseal.** Between the uninstall (08:38 UTC) and the verified restore (10:27 UTC) demo-hp's guest was destroyed, the agent and its pveum identity removed, the host re-enrolled, a new guest provisioned, the box re-claimed, the sealed off-site package opened with the customer's recovery code, an app redeployed and 3.8 MB restored. **Events recorded for demo-hp in that window: ZERO. Notifications: ZERO.** **Positive control on the query** (standing rule 3): the hub recorded **2 events all day across all customers**, newest `db_dump_completed` at 00:30:07 — so the store is reachable and the silence is real, not a bad filter. **The good half, stated first:** no FALSE alarm fired during a legitimate reinstall, which is what P7 was watching for. **The owed half:** `escrow_blob_served` exists precisely as the tripwire for this moment — its text is *"the blob cannot be opened without the customer's recovery code… If no recovery is in progress on that box, investigate"* — and it **has fired for demo-hp before** (twice, last 2026-08-04 20:12:54). Today's unseal, through the R-193 recovery screen, fired it **not at all**. Either the screen's unlock path does not emit it or the rebuilt-box path bypasses it; **which of those is not established here.** A reinstall and a theft of a machine look identical to the operator | **READY (M) — NEW 2026-08-09** | — | Emit on the recovery-screen unlock path; and decide which reinstall milestones are worth one line each | CC |
|
||||
| **R-281** | ~~**The hub said NOTHING through an entire reinstall — and the tripwire for a sealed-backup unseal did not fire on a real unseal.**~~ **WITHDRAWN 2026-08-09 — THE FINDING WAS AN ARTEFACT OF MY OWN MEASUREMENT, AND IT WAS WRONG IN BOTH DIRECTIONS.** The operator's mailbox settled it: the hub fired **twenty events** on 2026-08-09, and `escrow_blob_served` **DID** fire — 10:19:41 UTC / **12:19 CEST**, eight minutes before the verified restore. **Cause of the false reading, ESTABLISHED (not guessed):** the P7 query copied `/data/hub.db` **without `hub.db-wal`**. The hub runs SQLite in WAL mode (R-172), so every write since the last checkpoint was invisible. **The signature is an exact match:** P7 reported *"2 events all day, newest `db_dump_completed` 00:30:07"*, and the number of rows on 08-09 at or before 00:30:07 is **exactly 2**. **The two obvious alternatives were TESTED AND REFUTED**, not waved away: a **timezone offset** — all nine mailbox stamps equal the hub's UTC + 2 h exactly (`escrow_blob_served` 10:19→12:19, `host_down` 09:28→11:28, and seven more), so the window was right; and a **wrong customer key or wrong store** — the same table and key return the correct rows now. A live re-run cannot reproduce the fault because the WAL has since been checkpointed; the case rests on the command text plus the 2-of-2 count signature, and that is stated rather than dressed up as a reproduction. **This is a trap this project has already documented** — `operations/nodes.md` says copying `hub.db` alone is *"valid but stale … the worst failure shape"* — and I had avoided it correctly earlier in the same session before hitting it. **Split out: → R-285** (the real, opposite defect) and **→ R-286** (the measurement lesson) | **WITHDRAWN 2026-08-09** | — | Superseded by R-285/R-286 | CC |
|
||||
| **R-282** | **One secret, three different Hungarian names, and the email sends the customer to a page their box is not showing.** Sending it from the hub is „**Visszaállító** kód küldése"; the email that arrives is subject „Jelszó-**visszaállítási** kód", body „**Visszaállító** kód: …", and it instructs *„Add meg a vezérlőpult »**Elfelejtett jelszó**« oldalán"*; the page the box actually serves is „A szerver **beállítása**" asking for a „**Beállító** kód". **A rebuilt box shows a SETUP page and the hub can only send a RESET mail** (because hub-side the customer is still `claimed_at 2026-07-21`), so the instruction names a route that does not exist on screen. **It does work if you ignore the instructions** — the reset code was accepted on the setup page (302 + session), so this is naming, not function. **It cost this session real time and one wasted code:** the operator supplied a 3-word Hungarian code believing it was the recovery code, because the hub calls the claim code „Visszaállító kód" and the ESCROW code is also „Visszaállító kód" — the only reliable discriminator is length (claim = 3 Hungarian words; recovery = **10** EFF-list words, and the recovery screen does say „(tíz szó)") | **READY (S) — NEW 2026-08-09** | — | Pick one name per secret and use it on all three surfaces; make the mail's page reference match what a rebuilt box actually shows | CC |
|
||||
| **R-283** | **After a rebuild the hub says "Claimed 18d ago" while the box serves its first-run setup page.** `customer_claims` for demo-hp still read `claimed_at 2026-07-21 16:29:25`, `generation 2`, `issued_at 2026-08-03` while the freshly provisioned guest — whose `settings.json` is new — correctly showed „A szerver beállítása". The two sides never reconcile: the hub's claim state survives a guest rebuild and the box's does not. Consequences: the operator's screen says the box is claimed when it is not, a resend produces a RESET code instead of a SETUP code (→ **R-282**), and any previously issued code fails with *„Hibás vagy lejárt kód"* — a message that is technically true and tells the customer nothing about the real cause, namely their own reinstall. Mirror image of **R-214/R-235** (an already-paired box still told to pair itself) | **READY (S) — NEW 2026-08-09** | — | Let a report from a box carrying no claim state clear the hub's, or show both sides on the operator page | CC |
|
||||
| **R-284** | **„A kiválasztott tárhely majdnem megtelt." on a store that is 93 % FREE — an apparent inverted threshold.** Calibre-Web's deploy page rendered `<option value="/mnt/sys_drive" data-free-percent="93">` alongside „Tárhely (sys_drive) — **64.2 GB szabad**" and the warning „A kiválasztott tárhely majdnem megtelt." 93 % free read as 93 % used is the obvious candidate, and `checkStorageSpace(this)` is the function to look at. **Not confirmed by reading the code** — reported as measured output only. A capacity warning that cries wolf on an empty disk is one a customer learns to click past | **READY (XS) — NEW 2026-08-09** | — | Check `checkStorageSpace`'s comparison against `data-free-percent`; add a render test per branch | CC |
|
||||
| **R-285** | **A planned, supervised reinstall pages the operator as if the machine had died — there is no notion of expected downtime anywhere.** During the 2026-08-09 rehearsal the hub sent, all `status: sent` to the operator channel: `host_stale` 08:58 UTC, `node_stale` 09:00, **`host_down` 09:28 (error)**, **`node_down` 09:30 (error)**, `host_leaf_changed` 09:31, `host_recovered` 09:31, `node_recovered` 09:34, `offsite_delivery_stuck` 09:34 — eight operator mails for work that was deliberate, attended and announced. **This is the OPPOSITE gap from the one R-281 filed:** the alarms are not missing, they are indiscriminate. `host_stale` at 30 min and `host_down` at 60 min (`monitor/host_staleness.go:22-23`, `downAfter = 2 * threshold`) cannot distinguish a wiped-on-purpose box from a dead one, and `host_leaf_changed` firing on a reinstall is correct-but-expected. **Note the interaction with the mute used on 2026-08-09 evening:** blocking a customer silences everything, so today the only two settings are *page me for planned work* and *tell me nothing at all*. **What is owed is a middle:** a maintenance window, or an operator-set expected-downtime flag, that suppresses staleness and leaf-change while leaving genuine faults audible | **READY (M) — NEW 2026-08-09** | — | The evidence is the operator's mailbox plus `events`/`notification_log` for 2026-08-09 | CC |
|
||||
| **R-286** | **A control drawn from the same channel as the measurement cannot detect a defect in that channel — and this one passed while the measurement was wrong.** The P7 check asked *"did the hub record anything?"* against a stale snapshot, got "no", and then validated itself with *"is the hub recording ANY events today, for anyone?"* — **against the same stale snapshot**. It answered "2 events all day", which was internally consistent and entirely false. The standing rule (*an absent log line is not evidence*) was followed in form: a positive control WAS run. **It was the wrong kind of control**, and nothing in the rule as written says so. **The independent channel existed and was available the whole time: the operator's mailbox.** One glance at it would have shown eight alarms in the window. **The durable lesson, to be added where the standing rules live:** a control must come from a DIFFERENT channel than the measurement — same query, same snapshot, same API, same clock all fail this. **Concrete follow-through owed:** (a) add this to the standing rules in `runbooks/workspace-CLAUDE.md`; (b) any hub-state check in a runbook must copy `-wal` or query the pod directly, never `cat hub.db` alone — the trap `operations/nodes.md` already documents | **READY (S) — NEW 2026-08-09** | — | Parent: R-281 (withdrawn) | CC |
|
||||
| **R-287** | **`felhom-agent` CI is red for a TRUE reason, and the diagnosis it was filed under is wrong in every particular.** The task premise was *"the gate is sensitive to being run against a tag ref rather than a branch"*. **It is not.** `check-published-versions.py` enumerates releases from the **Gitea tags API** (`/api/v1/repos/admin/felhom-agent/tags?limit=200`, `main()`), so the checked-out ref is irrelevant; and the two previous tag pushes **passed** (run 190 `v0.126.0`, run 216 `v0.127.0`). **What is actually true:** run 267 (main, `28ba8593b8`, 2026-08-08 14:29 UTC) printed `ok v0.120.0: binary downloadable`; run 284 (tag, **the same commit**, 2026-08-09 09:30 UTC) printed `FAIL v0.120.0 — binary NOT downloadable (HTTP 404)`. **A published release became uninstallable between those two runs.** The registry now holds exactly the ten newest versions (0.121.0…0.128.0); `0.128.0` was published **2026-08-08 16:47 CEST = 14:47 UTC, eighteen minutes after run 267**, and `0.120.0` is gone. **WHO REMOVED IT IS NOT ESTABLISHED, and that is stated rather than guessed:** `package_cleanup_rule` is **empty** (queried in Postgres), `app.ini` sets no package limit, `publish-agent.sh` only pre-deletes the version it is publishing (`:77`), the Gitea pod has **53 days uptime and 0 restarts** so `RUN_AT_START` did not fire, and **no `DELETE` on the packages API appears in 48 h of Gitea router logs**. The leading candidate is the internal `[cron.cleanup_packages]` `@midnight` job, which falls inside the window and would leave no router log line — **leading candidate is not established.** **THEREFORE NO GATE WAS SILENCED AND NO WORKFLOW WAS CHANGED.** Silencing it would hide a released-but-uninstallable version, which is the exact R-115 defect the gate exists to catch. **It will recur:** if the ten-version window is real, the next publish evicts `0.121.0`. **Two honest fixes, both out of tonight's scope:** bound the gate to versions at or above the vouched `min_agent` floor (0.127.0 today — nothing installs 0.120.0 and nothing can), or retire ancient tags when their packages go. **Also measured, and good news:** the failure alarm DID send — `RESEND-ACCEPTED id=fa1a7a83-714f-4357-b0ca-d3c4bb7ae73f` | **READY (M) — NEW 2026-08-09** | — | Establish the deleter first; do not raise the retention until it is known | Viktor |
|
||||
| **R-288** | **The capability map is too long to be read, and that is why it stops being true.** `architecture/00-capability-map.md` is **134 642 bytes / 19 456 words across 99 table rows in only 159 lines** — because the rows ARE the length. Measured, longest first: the unaided-recovery-journey row is **3 024 words**, the offsite-password-recovery row **1 087**, the unattended-restore-proof row **971**, the app/guest-network-failure row **904**. That single longest row is a novella of nested corrections, each appended rather than resolved. Its own verification stamp reads **2026-07-16 against evidence corpus @ felhom.eu tip `4b18cc5`** (line 23) — three weeks stale, which is the measurable consequence: nobody re-reads a row they cannot finish. **This is the project's memory, so restructuring it is surgery and wants daylight** — filed, deliberately not attempted in the 2026-08-09 session. **What the shape should probably be:** one line of status per capability plus a dated evidence pointer, with the argument moved to the audit it came from | **READY (M) — NEW 2026-08-09** | — | Do not fold this into another session; it needs its own | Viktor |
|
||||
| **R-289** | **R-182's register row describes a defect the code no longer has — an OPEN row that is a false alarm.** The row reads *"A full disk tells the operator about ONE app and silently swallows every other app's refusal for an hour"*, cited at `hub/internal/notify/dispatcher.go:268`. **Read against live source 2026-08-09, that is fixed:** the per-run digest `backup_run_failures` is allowlisted (`hub/internal/api/handler.go:1837`), operator-only (`dispatcher.go:423`) and templated (`notify/templates.go:48`); `recovery_unit_capture_failed` is now a **record-only** event (`dispatcher.go:376`) whose notification IS the digest, listing every failed app in one mail; and a cooldown drop now writes a `suppressed` row instead of vanishing (`dispatcher.go:314-330`). The capability map already records the fixed shape (*"EVERY failing app, in ONE mail per run"*). **So the register is behind the code, which is the mirror of the decay this session was looking for** — the session expected stale PROOFS and found a stale DEFECT. **Not closed here, deliberately:** the digest's *delivery* has never been observed end to end (the page's own "an app crashes — the email leg has never been confirmed" card), so the honest move is to re-scope R-182 to that residue rather than tick it | **READY (XS) — NEW 2026-08-09** | — | Re-scope R-182 to "the digest has never been seen delivering", or close it and open that | CC |
|
||||
| **R-290** | **Most capability-map rows that back a green dot cite no evidence document at all — measured, 20 of 28 probed.** The page's *Walked* means *"done end to end on real hardware, evidence on file"*. Extracting the evidence column for the 28 rows behind the page's claims found a `tests/` or `audits/` path in **8**; the other 20 carry prose only. **Consequence, applied this session:** of 32 claims the page drew as Walked, **12 were downgraded to Built** because no walk document exists for them — `install.installer-by-tag`, `use.lifecycle`, `drives.enrol`, `drives.migrate`, `backup.tier1`, `backup.whole-machine`, `backup.restore-proof`, `fault.selfheal`, `fault.operator-email`, `fail.drive-filling`, `fail.lost-recovery-code`, `fail.hub-down`. **This is not a claim that those twelve are false** — several are near-certainly fine — it is a claim that nothing on file distinguishes them from an opinion, which is exactly what the status word promises. **The gate now enforces it going forward:** `scripts/check_stands.py` fails on `status: walked` with no `evidence:` source. **What is owed:** either a walk document per row, or an honest demotion in the map itself (the map is the source; the dataset only follows it) | **READY (M) — NEW 2026-08-09** | R-288 | The dataset was corrected; **the capability map itself still says PROVEN-LIVE for these rows** and is the thing to fix | Viktor |
|
||||
|
||||
**Explicitly still open, untouched by this session:** R-246, R-255, R-256, R-257, R-261, R-262,
|
||||
R-263, **R-264** (the twenty-one undecided facts — a design session of its own), R-240, R-243,
|
||||
|
||||
Reference in New Issue
Block a user