diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index f13f146f..f3ae0d6e 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -165,3 +165,40 @@ | **R-329** | **`app_start_failed` was emitted with severity `"warn"`, so every one of them was delivered to nobody.** Shipped in controller v0.223.0 (+ hub v0.107.0). Evidence: `audits/DRILL-r329-r386-2026-08-23/`. **Reasoning kept:** *The vocabulary is EXACT and it is the HUB's, not ours* — `{info, warning, error, critical}`; anything else is coerced to `info` at ingest and dropped by `severityNotifies` before BOTH legs. **This was the SECOND occurrence** (`DiskAlertKind.Severity` until v0.215.0), and its comment had recorded the lesson — **a comment is not a guard**, so the guard is now an AST walk over the whole controller, with the six variable-passing call sites registered by name because a walk cannot follow a variable and *an unlisted limit is not a limit, it is a hole*. **The register's own framing was that the DECISION was the work** — should a stopped app mail the customer at all? Answered: **operator always, customer OFF by default**, because `processOperator` never consults customer preferences, so one word fixed the operator leg and left the customer leg exactly where the ruling wanted it. **Deliberately NOT added to `operatorOnlyEvents`** — that would make the new toggle visible, flickable and structurally incapable of delivering. **Measured on the live hub DB: 91 events stored all-time, ZERO notification rows before the fix; one operator row, `warning`/`sent`, after it.** | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.223.0 + hub v0.107.0, 2026-08-23) | full text: `git show 68a9f5475cd2:documentation/backlog/OPEN-ITEMS.md` | | **R-386** | **A single-container app stopped out of band raised no alarm, and a comment stated the opposite as settled fact.** Shipped in controller v0.223.0. Evidence: `audits/DRILL-r329-r386-2026-08-23/`. **Reasoning kept:** *the state test was guessing at something the product already knows.* `DesiredState` records the customer's intent, has **exactly one writer**, and is tri-state; `StateExited` never survives aggregation, so no state test can separate an out-of-band stop from a customer stop. The ruling: `Stopped` → no alarm, `Running` → **alarm**, **absent → UNKNOWN, keep today's behaviour AND announce it**. *Reading unknown as "nobody asked" would, on the first cycle after upgrade, e-mail about every app any owner ever deliberately stopped — fleet-wide, from a field that predates the intent it is being asked about.* **A rule without a mechanism is a wish:** every such suppression sets `IntentUnknown` and the names are logged at INFO, so an operator can answer *"how many apps am I blind to?"*. **`failedRestart` must still lift a `Stopped` intent or F-CRIT-1 re-opens.** **Fenced act: adding a `DesiredState` WRITER** — twelve of `StopStack`'s fourteen callers are machines. Proven live: alarm 24 s after an out-of-band `docker compose stop`, heartbeat `1 currently down` against the previous day's `0`; and with intent removed, suppressed *plus* the log line naming the app. **0 of 8 deployed apps on `demo-hp` carry an absent intent.** | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.223.0, 2026-08-23) | full text: `git show 68a9f5475cd2:documentation/backlog/OPEN-ITEMS.md` | | **R-389** | **Only the FIRST broken app per hour reached the operator — the cooldown key named the event type, not the app.** Shipped in hub v0.108.0. Evidence: `audits/DRILL-cooldown-grain-2026-08-23/`. **Reasoning kept:** the fix is a THIRD SIBLING of `cooldownTierSuffix`/`cooldownRunSuffix`, separate for the reason the second one's docstring already gives — *the existing two keep byte-identical semantics for every type that uses them.* **`cooldownStackSuffix` takes the EVENT TYPE as well as the details, unlike its siblings, and that asymmetry is the whole safety property:** `tier` and `run_id` appear only on types that want that grain, `stack_name` does not. **`perAppCooldownEvents` is a named allow-list with `app_start_failed` and nothing else** — *the backup family's cooldown is coarse ON PURPOSE (R-97a, R-182) so one full disk sends one digest rather than one mail per app*, and this is not hypothetical: **`crossdrive_failed` is severity `error`, reaches the operator leg, and carries `stack_name` through a different struct**, so a payload-shape rule would have split it silently. **The fenced act is adding an entry for a type whose family has a digest or a coarse-by-design cooldown.** `app_start_failed` qualifies precisely because it has NO digest — there is no `apps_down_run` the way `backup_run_failures` summarises a run. **The hour is unchanged; the grain was the complaint.** Fail-soft: absent or malformed details degrade to the old key and the mail still goes. **PROVEN LIVE 2026-08-23:** two apps four minutes apart gave **2 sent / 0 suppressed** where the same shape gave 1 and 1 the day before, each repeat suppressed under its OWN key (`…:opengist`, `…:calibre-web`) against the previous day's shared `key=demo-hp:app_start_failed`; and `crossdrive_failed` for two different apps stayed **coarse** under `key=demo-hp:crossdrive_failed`, byte-identical to the derived v0.107.0 value. **AND IT WAS NEVER FILED UNTIL THE DAY IT WAS FIXED** — it lived in a REPORT.md observations paragraph, which is why gate 11 now exists. | **CLOSED — SHIPPED + PROVEN-LIVE** (hub v0.108.0, 2026-08-23) | full text: `git show 45659bdc5a2f:documentation/backlog/OPEN-ITEMS.md` | + + +## 2026-08-30 — the restore tells the truth (controller v0.226.0) + R-395 + +Six rows closed. **Full original text: `git show e027b5d9 -- documentation/backlog/OPEN-ITEMS.md`.** +Compressed here to title, shipping version, evidence, and the sentences that state a RULE. + +| ID | Title | Shipped | Evidence | +|---|---|---|---| +| **R-353** | A local unit restore reported a bare completion whether it returned an entire dataset or nothing | controller v0.226.0 | `documentation/audits/evidence-r353-r360-live-2026-08-30/live-validation.txt` — live sentence `A(z) opengist: 1 adatkötet visszaállítva — az alkalmazás újraindult.` | +| **R-357** | The destructive reconstitute had no free-space gate; all three that existed guarded non-destructive paths | controller v0.226.0 | seam tests only — **NOT live-validated**, by design | +| **R-358** | `OffboxFullScratchReady` asked "non-empty directory", which is what a failed restic run leaves | controller v0.226.0 | `documentation/audits/evidence-r353-r360-live-2026-08-30/live-validation.txt` — marker `{"schema":1,…,"full":false}`, gate logged place-to-live closed | +| **R-360** | The verification-copy delete gated on the concurrency flag, which a verification restore never holds | controller v0.226.0 | `documentation/audits/evidence-r353-r360-live-2026-08-30/live-validation.txt` — refused in the live flag state; planted canary survived | +| **R-396** | A unit-only verification restore unlocked the DESTRUCTIVE full restore | controller v0.226.0 | same evidence; found while answering R-358's open question | +| **R-395** | `STATUS.md` contradicted itself about the controller version | doc fix, same session | `STATUS.md` at `e027b5d9` | + +**The rules these leave behind — the reason the rows are kept rather than deleted:** + +- **A restore outcome is a claim about THE BACKUP, never about the app.** R-355 extended to the Tier-1 + path. The off-site twin has `SafetyDump` as an honest discriminator; the local path has none, so no + claim about the app is available to it at all. Not merely unproven — unprovable from a manifest: + §6.3 records that an absent dump has causes that say nothing about the app. +- **Zero-replayed has two causes and they are opposite news.** "The backup held no data" and "the + backup listed data that did not come back" must never share a sentence. +- **The destructive reconstitute uses NO headroom margin**, matching `PlaceOffsiteRestore`. The ×1.1 + elsewhere exists because that gate PREDICTS a download; this one measures a tree that already exists. +- **Fail closed when a probe reads ≤ 0.** `free < need` with `need == 0` is FALSE, so an unmeasurable + input sails through — a gate present and inert, which is worse than no gate because it reads as + protection. +- **A hidden button is not a guard.** Template enable-flags control a button; the handler must refuse. +- **One boolean must not drive three intents** (R-396): `ScratchReady` answered "is there a scratch" + while being consumed as "may we place" and "may we destructively restore". +- **Never restate a version in a second place on the same page** (R-395). Live versions belong in the + hub, never in a doc. +- **A doc comment claiming a guard exists is why nobody looks for the missing guard** (R-360). Correct + such a sentence in place; do not delete it. + diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 692d2343..d81dc28e 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -136,6 +136,8 @@ the fault was real. Full observables: `tests/campaign11-evidence-2026-08-05/jour | ID | What | State | |---|---|---| +| **R-397** | **`NotifyIntegrityOK` and `NotifyIntegrityFailed` are dead code — the controller runs no integrity check at all.** Both exist in `internal/notify/notifier.go` and a repo-wide grep finds **no caller**. The `backup_integrity_ok` / `backup_integrity_failed` event types are therefore unreachable, and the hub's customer Backup card rendered `integrity_ok` from a field nothing writes until v0.109.0 dropped the row (R-331). There is also a `monitoring.ping_uuids.backup_integrity` config field and a „Mentés integritás — Hetente (vasárnap)" line on the controller's own monitoring page, **so the product currently tells the operator an integrity check runs weekly and no such check exists.** Noticed twice in one day, from opposite directions (R-331's dead report fields, then this task's restore work), which is why it is a row rather than a note. | **OPEN — MEDIUM** | — | **Decide first WHETHER an integrity check is wanted, then delete or build — do not leave the third state.** If wanted, `restic check` is R-359's territory and these notifiers are its natural sink; if not, remove the notifiers, the event types, the ping UUID and the schedule line together, because the schedule line is the part that actively misleads. | CC | +| **R-398** | **`resticStep` is not a seam, so no test can drive any restic-backed path.** Every off-site operation funnels through it, and it shells out — so `RestoreOffboxScratch`, `PlaceOffsiteRestore` and the whole capture side can only be unit-tested up to the point restic would run. Felt directly in v0.226.0: R-358's safety property is an ORDER (clear the marker before restic, write it after), and with no seam that order could only be pinned by an **AST walk** of the function rather than by executing it. That works and is honest about what it proves, but it is a structural workaround for a missing seam and would not catch a reordering introduced through a helper. **Contrast, and it is the argument for the row:** `offboxLatestSnapshot` gained a seam in this same release (`SetOffboxLatestSnapshotFn`) precisely because a correctness gate could not otherwise be proven, and that one took four lines. | **OPEN — SMALL** | — | Add `resticStepFn` beside the existing `offboxFreeFn` / `offboxSizer` / `offboxLatestSnapFn` seams, nil in production, and convert R-358's AST test to an execution test. **Do it as its own change, not folded into a bugfix** — a seam added under pressure is how a test ends up asserting the shape of the thing it was written beside. | CC | | **R-385** | **A controller was built, baked AND vouched with no CHANGELOG entry of its own, and every gate stayed green.** Controller **0.221.1** shipped on 2026-08-23 while the newest heading in `felhom-controller/CHANGELOG.md` still read `v0.221.0` — the prune-ordering fix (commit `810b18a`) had been written INSIDE the v0.221.0 entry instead of getting its own. The image was never in question; the RECORD was, and the fleet ran a version the record did not name. **`scripts/golden_currency_gate.py` could not catch it by construction:** it failed only on `released > baked`, so a golden AHEAD of the record passed silently. Measured on the real history: `newest released 0.221.0 / newest golden baked 0.221.1 → OK, exit 0`. | **CLOSED — 2026-08-23** | — | **Both halves fixed, both directions red-proofed.** The record: `v0.221.1` has its own heading carrying the MOVED (not duplicated, not deleted) reasoning — commit `da75603`, pushed alone before anything else. The gate now asks *"is the baked version WRITTEN DOWN?"* — the baked version must have its own `## vX.Y.Z` heading **anywhere** in the CHANGELOG. **Membership, not `baked > released`, deliberately:** a comparison against the newest heading alone goes green the moment any later entry is written, leaving the unrecorded version permanently unrecorded and the gate permanently silent about it. INCONCLUSIVE (exit 2) preserved. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/evidence/gate-0*.txt` — old gate/old record `exit 0`, new gate/old record `exit 1`, new gate/fixed record `exit 0`. | CC | | **R-387** | **The hub REWRITES an unknown severity and says nothing, and the guard built to catch that sits downstream of the rewrite.** One handler, two fields, opposite discipline: an unknown `event_type` is rejected with a loud `400`, while an unknown `severity` was silently coerced to `info` — after which `severityNotifies` drops it and NEITHER delivery leg runs. **Two shipped features went out that way**: `DiskAlertKind.Severity` emitted `"warn"` until controller v0.215.0, `app_start_failed` until v0.223.0. **Measured on the live hub DB 2026-08-23: 91 `app_start_failed` events stored all-time and ZERO `notification_log` rows before that day** — not one, on any channel, while every POST returned 200. **The dispatcher's `unrecognized severity` line could never execute** for an API event, because the coercion one line upstream guarantees the value it looks for cannot arrive. | **CLOSED — hub v0.107.0, 2026-08-23** | — | **The coercion STAYS; only the silence is fixed** — a rejected event is a LOST event, and losing an alarm is worse than mis-routing one. A `WARN` now names the customer, the event type, the rejected value and the consequence. **The dispatcher branch was KEPT, on evidence not caution:** `cmd/hub/main.go` wires `dispatcher.ProcessEvent` DIRECTLY as the `monitor.EventNotifyFunc` for the staleness, host-staleness and offsite-box checkers, which never pass through the handler — for them it is the only severity guard there is; deleting it as "dead" would have removed the live half while the dead half supplied the justification. All 90 severity literals in `internal/monitor` verified already valid. Proven live: `[WARN] [api] Event from demo-hp: severity "warn" is not in {info,warning,error,critical}…`, with an `error` control silent. Evidence: `audits/DRILL-r329-r386-2026-08-23/evidence/live-19-scenarioH-after.txt`. | CC | | **R-391** | **Gate 11 (observations) is registered in three of the four runners; `app-catalog-felhom.eu` is the exception.** The controller and agent runners already carried a shared-gate mechanism (`SHARED_REUSE`, `SHARED_INSTRUCTIONS` pointing into `felhom.eu/scripts/`), so registering there was one constant and one `GATES` line each. **`catalog_gates.py` has no such mechanism:** its `run_gate` joins every entry against its OWN `scripts/` directory, so it cannot invoke a sibling repo's script at all; and its loop appends `--all` to every gate unconditionally, which the observations gate would read as a path. Registering there therefore needs `run_gate`'s contract widened AND the argument handling changed — a refactor of a runner whose shape is deliberately different (per-app scoping, network/runtime gates excluded from `--fast`), in a repo this task marked out of scope. **The exposure today is nil** — `app-catalog-felhom.eu/REPORT.md` has no observations section, and the gate passes quietly on that — but a future catalog session could write one and nothing would read it. **Filed rather than left as a sentence in a report, which is the exact failure R-389 records.** | **OPEN — LOW** | — | Either give `catalog_gates.py` the `SHARED_*` absolute-path mechanism the other two runners already have and stop appending `--all` to gates that do not take it, or state in that repo's CLAUDE.md that its REPORT.md carries no observations section by convention. **Do not copy the gate script** — the shared checker lives in ONE place (`felhom.eu/scripts/`) and copying it is the drift the shared pattern exists to prevent. | CC | @@ -531,13 +533,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-348** | **Every agent restart blanks the reported backup list for up to ~18 hours, and the comment that covers it says "unaffected".** Observed 2026-08-20 while deploying R-344: the first host reports after `demo-hp`'s agent restart carry **`0 backups`** (11:15:50 and 11:30:52 CEST, two consecutive), while the box's own `pvesm list` shows archives present on **both** tiers. `internal/backup/store.go`'s `Store` is in-memory and `byTarget` is repopulated only when a backup **runs** — daily for the local tier, weekly for offsite — so the field reads 0 until the next run. `restore_tests` did **not** blank, because that half has a durable on-disk companion (`RestoreTestState`, R-189). **It blinds no alarm, and that was CHECKED rather than assumed.** `hub/internal/monitor/deadline.go` scans back over stored reports with a 7-day `backupEvidenceLookback` whose own comment names this exact case — *"when the LATEST report carries none... and against an agent that stayed restarted for days"* — and `pbs_snapshots` stayed populated at 2 regardless. So this is an observability wart, **not** a safety hole, and it is filed at that severity deliberately. **What is actually wrong is the comment.** The `Store` doc says *"Backups are unaffected — their freshness has a ground truth on the storage (R-84)"*. That is true of the **consequence** and false of the **field**, and it sits three lines below a paragraph explaining that the very same sentence about restore-tests *"used to be here and it is now FALSE"* — so the file already carries one correction of this shape and invites the next reader to trust the surviving half. | **READY (XS) — NEW 2026-08-20** | — | Say what is measured: the field IS lost on restart and repopulates only when a backup runs; the freshness VERDICT is unaffected because the hub looks back 7 days. **Name `backupEvidenceLookback` in the comment** so the cross-repo dependency is visible from the agent side — today the agent's claim of safety rests on a hub constant it does not mention. Per the workspace rule, a comment asserting an invariant needs a test pinning it: the pin belongs on the HUB side, asserting the verdict survives a report carrying `backups: []`. | CC | | **R-349** | **"Prove it by hand, then publish" leaves the fleet running a DIFFERENT binary under the SAME version name — and self-update cannot notice.** Hit on 2026-08-20 during the R-344 train, caught and corrected the same hour, filed because the next prove-then-publish train will hit it identically. **The mechanism:** a proof deploy is a hand build (`go build -ldflags "-X main.version=0.130.0"`), while `scripts/release-agent.sh` deliberately builds with **`-trimpath -buildvcs=false`** so the published artifact is reproducible (R-186). Same source, same version string, **different bytes**: `256e0829...` on the boxes vs **`a56a92a7...`** published and vouched. **Nothing corrects it automatically**, and that is the sharp edge: the boxes already report `0.130.0`, so the self-update path sees the vouched version as already installed and does nothing, **forever**. The divergence is invisible to every version check in the system — the hub, `--version`, and the artifact manifest all agree, because they all compare the version STRING. **Consequence if unnoticed:** the binary a customer box runs is not the binary the operator vouched, and not the one a reinstall would fetch — so a bug reproduced on the fleet may not exist in the published artifact, or vice versa. It is the same "one version name, two binaries" hazard `publish-agent.sh` already carries a comment about for `CGO_ENABLED`; that comment fixed the two ENTRY POINTS and does not cover a hand build during a proof. **Corrected here** by downloading the published artifact from the registry (not rebuilding it locally — the boxes get the bytes a fresh install would get) and installing it on both; both now report `sha256 a56a92a7...`, matching the vouch. | **READY (S) — NEW 2026-08-20** | — | Make the reconciliation a step, not a memory: the honest fix is for the agent to REPORT the sha256 of its own binary in the host report, so the hub can compare it against the vouched `agent_sha256` and flag drift — **exactly the mechanism `wrapper_sha256` already implements for the PBS wrapper** (R-50b), whose manifest help text says it *"makes host drift visible: agents report the installed file's hash and a mismatch is surfaced on the host page"*. The pattern exists and is proven; it simply was never extended to the agent's own binary. Cheaper interim: end every prove-then-publish train by installing the DOWNLOADED artifact. | CC | | **R-350** | **SECURITY — the hub operator password was printed in cleartext into a session transcript by CC, 2026-08-20. Rotation recommended.** **What happened:** vouching the artifact manifest used `curl -w '%{redirect_url}'` for confirmation. The hub answers `POST /configuration/artifacts` with a **303**, and curl renders the redirect target **with the basic-auth credentials re-attached** — so the URL it printed contained `http://:@10.43.52.34:8080/configuration?flash=artifacts_set`. The password was never read aloud from the credentials file, never echoed deliberately, and every other call in the session correctly printed only `${#HUB_PW}`; it arrived through curl's own output formatting, which is why the usual discipline did not catch it. **Blast radius, stated precisely rather than minimised:** the value is **not** in git, not in `CHANGELOG.md`/`REPORT*.md`/any committed file (checked), and not in the evidence directory — it is in the Claude Code session transcript under `~/.claude/projects/` on DooPlex, which is operator-readable and persists across sessions. The hub UI is reachable only on the k3s ClusterIP and via the operator's own routes, not from the internet. **The value is deliberately not recorded here; it is stored out-of-band in the usual credentials file.** | **READY (S) — NEW 2026-08-20** | — | **Operator decides whether to rotate.** The hub's own `/configuration` password form does it (`current_password`/`new_password`/`confirm_password`), and per `hub-password-ui-2026-07-13` the DB override wins over the ConfigMap, which stays break-glass. CC can perform the rotation **file-to-file without printing the new value** (the `operator-present-one-time-secrets` convention) if asked — it did not do so unilaterally, because rotating a credential the operator holds in their own head or notes is their call, not CC's. **The reusable half, which matters more than this one password:** never use curl's `%{redirect_url}` (or `-v`, or `--libcurl`) against a basic-auth endpoint — all three re-render the credential. Confirm a redirect with `%{http_code}` and read the flash from a follow-up GET. | **Viktor decides**, CC executes | -| **R-395** | **`STATUS.md` contradicted itself about the controller version, on the page the operator reads first.** Filed by the task spec, verified 2026-08-30 at `c2c1fb4` and **fixed in the same session**. The „Waiting on you" block said the golden carried 0.223.0 and the floor read 0.222.0; „What works", fourteen lines below, said `demo-hp` ran **0.219.0** and the fleet floor was **0.218.0**, cross-referencing „item 2 above" — which said „Nothing else". The second pair was stale from the 0.218/0.219 release and had survived at least two edits of the block above it. **A status page that disagrees with itself is worse than one that is merely out of date**: the reader cannot tell which half to trust, so they trust neither and go and look, which is the work the page exists to save. **The structural cause, and it is the reusable part:** the same fact (what version is live) was written in two places with no link between them, and only one had a reason to be touched during a release. The fix removes the duplicate rather than correcting it — „What works" now points at the item above instead of restating a version. Consistent with the standing rule that live versions belong in the hub, never in a doc. | **CLOSED — fixed 2026-08-30, same session** | — | Done. The remaining discipline is the rule, not the row: never restate a version in a second place on the same page. | CC | -| **R-396** | **A unit-only verification restore unlocked the DESTRUCTIVE full restore, and it is reachable by the most ordinary customer action.** Found while answering the R-358 task spec's open question, which asked whether the flow can reach that state; it can, and by the safest-looking route on the page. **The chain, each link read from source:** „Ellenőrző visszaállítás" (`mode=unit`, advertised as non-destructive and the default) calls `RestoreOffboxScratch(ctx, app, full=false)`; `offboxRestoreScratchDir` **ignores `full`**, so both modes write the SAME directory, and restic's `--include` limits WHAT is extracted, never WHERE; the wizard sets `ScratchReady` from `OffboxFullScratchReady`, which pre-fix answered „the directory exists and is non-empty"; `deriveWizardStep` then derives **both** `PlaceEnabled` and `RestoreEnabled` from that one flag. So a customer who ran the SAFE restore was afterwards offered „Teljes visszaállítás indítása" over a copy holding only the recovery unit. **This is strictly worse than R-358 as filed**, which assumed the bad state needed a FAILED download; it needs only a successful safe one. | **CLOSED — controller v0.226.0 (2026-08-30)** | — | Closed by R-358's marker, which records `full` — a unit-only scratch reads not-ready and the card stays shut. **PROVEN LIVE on demo-hp 2026-08-30**: after a `mode=unit` restore the marker read `"full":false` and the gate logged „scratch holds a UNIT-ONLY restore … place-to-live stays closed". Pinned by `TestR358_UnitOnlyScratchClosesTheFullRestoreCard`, which asserts the WIZARD view (place/restore closed, prepare open, verify still available) rather than the predicate alone. **The lesson to carry: one boolean drove three different intents.** `ScratchReady` answered „is there a scratch" while being consumed as „may we place" and „may we destructively restore" — three questions, one flag, and the weakest of the three set the answer. | CC | -| **R-353** | **A restore reported success having returned configuration and no data — and no screen could have told the customer.** `demo-hp`, 2026-08-21, OpenGist. The off-site reconstitution refused at 16:37:14 (not installed); the person reinstalled and ran the local unit restore, which reported `Restore-from-unit completed: opengist in 8.666896042s`. **The unit it restored from contains `manifest.json` + `compose/{app.yaml,.felhom.yml,docker-compose.yml}` and NOTHING else — `volume_dumps: None`, `db_dumps: None`** — and the off-site snapshot was **182.3 KB**. So the restore returned the app's configuration; there was no data leg in the unit to return, and the outcome said only that it had completed. **A warning beside a success is read as a success, and an unknown must never be drawn as healthy.** **Compounding, and recorded as UNKNOWN rather than fine:** whether the 40-class reaches the off-site tier at all has **not been observed** — `runVolumeDumps` (`backup/backup.go:607+`) covers them on paper, but every unit on the box reported `volume_dumps: None`, including `calibre-web` on the data drive, because no nightly dump run had happened on a one-hour-old box. **CLOSED 2026-08-30, controller v0.226.0.** `RestoreFromRecoveryUnit` now returns `(UnitRestoreResult, error)` and `unitRestoreOutcomeMsg` turns it into the sentence. **The count was never missing** — `restoreDockerVolumesFrom` always returned it and the one-line wrapper discarded it, so the surface was structurally unable to say what came back. THREE cases, because zero-replayed has two causes that are opposite news: data returned (named and counted); nothing returned and the unit listed nothing („ez a mentés csak a beállításokat tartalmazta, adatot nem"); nothing returned though the unit listed dumps („…de egyik sem állt vissza. Az adataid változatlanok maradtak."). **Every sentence is a claim about THE BACKUP, never about the app** — this path has no `SafetyDump` discriminator and §6.3 records that an absent dump says nothing about the app (R-361 destroyed canonical `.sql` files for four months). R-355 extended to Tier-1; recorded in the controller's CONTEXT.md. **PROVEN LIVE on demo-hp 2026-08-30**, through the endpoint the UI invokes, read off the customer's own wizard page: `A(z) opengist: 1 adatkötet visszaállítva — az alkalmazás újraindult.` (1 volume of 1 listed, 0 databases of 0 listed — correctly no database clause). **Scenario B could NOT be reproduced live, and that is stated rather than glossed:** no app on `demo-hp` still has a data-less unit — the drill's opengist has been recaptured and now lists one volume dump. Falsifying a manifest to produce it would be the hand-set-state shortcut this project forbids, so that branch is carried by `TestUnitRestoreOutcome_BackupHeldOnlySettings` and the A5 seam test. Evidence: `documentation/audits/evidence-r353-r360-live-2026-08-30/live-validation.txt`. | **CLOSED — controller v0.226.0 (2026-08-30). Leg (1) shipped; leg (2) was satisfied by the 2026-08-21 drill** | — | **Two things, in order. (1)** A restore whose unit carries no `db_dumps` and no `volume_dumps` must **say so in its outcome** — „a mentés csak a beállításokat tartalmazta, adatot nem" — instead of reporting a bare completion. The verdict must consult what was actually placed, not merely that the operation ended. **(2)** Then *prove* the off-site coverage of a named-volume app by running a dump cycle and reading the resulting manifest, rather than inferring it from the gate order. Do not close (1) on the strength of (2) being likely. **(2) IS NOW SATISFIED — drill 2026-08-21.** A dump cycle was run and the manifests read: `privatebin volume_dumps=[privatebin_privatebin_data.tar]`, `opengist volume_dumps=[opengist_opengist_data.tar]`, `kimai volume_dumps=[kimai_kimai_db_data.tar, kimai_kimai_var.tar]` — the 40-class DOES reach the off-site tier, and PrivateBin's planted 1 MB came back byte-identical from its off-site snapshot into the checking folder. **(1) stands and is now strictly larger than when written:** R-354 shows the bare completion is also reported over a unit that DID carry a data leg, because the off-site restore never replays volume dumps at all. | CC | -| **R-357** | **The DESTRUCTIVE restore has no free-space gate; the three that exist are all on non-destructive paths.** `offbox_reconstitute.go` contains **zero** references to `offboxFree`; the gates sit at `offbox_restore.go:231` (scratch restore), `:297` (prepare-full) and `:423` (place-to-live). Proven 2026-08-21 23:11 with 300 KB free and 1 MB to write: it stopped `paperless-ngx`, failed halfway (`rsync … No space left on device (28)`), left the data directory holding **2 of 5** planted entries, and restarted the app. The message is honest but is raw rsync output. **CLOSED 2026-08-30, controller v0.226.0.** The gate sits before `mapOffsiteRestorePaths`, before `writeSafetyDump` and well before `StopStack`, so a refusal costs the customer nothing — **position is the whole fix**, and `TestR357_DestructiveRestoreRefusesWithoutHeadroom` asserts the StopStack call count is 0 rather than the error text, because a gate placed after the stop returns the right sentence and still takes the outage. Wording shared with the three sibling gates via `offsiteNoSpaceMsgFmt`. **No headroom multiplier** — this copies a measured tree; `OffboxRestorePrepareFull`'s ×1.1 predicts a download (ruling in the controller's CONTEXT.md). **Fail-closed when either probe reads ≤ 0**, a real hole and not a formality: `free < need` with `need == 0` is FALSE, so an unmeasurable scratch sailed straight through — a gate present and inert. **NOT live-validated, deliberately:** filling a real filesystem is a drill step, not a build step; the seam tests carry it (`SetOffboxFreeFn`, `SetOffboxSizer`, and the new `SetOffboxLatestSnapshotFn`). **The red-proof exposed a hollow test and it is recorded rather than quietly fixed:** the first fixture refused earlier at the placement stat pre-pass, so `stops == 0` passed against the pre-fix code. Corrected, and the assertions reordered so a removed gate now reports `THE APP WAS STOPPED (1 call(s)) … (err=)`. | **CLOSED — controller v0.226.0 (2026-08-30)** | — | Same gate, same wording as `:297`, before the stop. | CC | -| **R-358** | **A FAILED scratch restore leaves a partial copy that the product then offers as a full restore source — and the destructive restore runs from it and reports success.** `OffboxFullScratchReady` (`offbox_restore.go:305`) asks only whether the directory exists and is non-empty; its comment defers completeness to `PlaceOffsiteRestore`, which stats top-level placements, not files. Proven 2026-08-21 22:54-22:56 against a deliberately corrupted store: the restore failed honestly (`ciphertext verification failed`, 54 files, 15 of 16 originals), the wizard then offered „Teljes visszaállítás indítása", and pressing it reported `ok=true`. **The failure is detected and then forgotten.** **CLOSED 2026-08-30, controller v0.226.0.** `RestoreOffboxScratch` writes `.felhom-restore-complete.json` (0600, tmp+fsync+rename) **only after restic returns nil**, and clears any stale marker **before** restic starts; `OffboxFullScratchReady` reads it. Absent, unreadable, wrong schema or `full:false` → not ready, with a WARN naming which. **Both handlers refuse server-side** — the wizard flags control a button, and a hidden button is not a guard; a direct POST over a part-copy used to be accepted. Both orderings are pinned by an AST test, because `resticStep` is not a seam and the order IS the safety property. **PROVEN LIVE on demo-hp 2026-08-30**: a `mode=unit` restore wrote `{"schema":1,"snapshot_id":"84542ec8","full":false,…}` at mode 600 with no `.tmp` left, and the gate logged `kimai: scratch holds a UNIT-ONLY restore (snapshot 84542ec8) — not a full copy, so place-to-live stays closed`. Evidence: `documentation/audits/evidence-r353-r360-live-2026-08-30/live-validation.txt`. **See R-396 for what the investigation into this turned up.** | **CLOSED — controller v0.226.0 (2026-08-30)** | — | Record the failure against the scratch and refuse to place from it until it is re-prepared. | CC | | **R-359** | **The off-site restic store is never verified by anything, ever.** The complete set of restic verbs in the controller is `restore, snapshots, backup, unlock, stats, init, forget, prune, cat` — **no `check`**. The agent's `RestoreTest` is PBS-tier only. Established 2026-08-21 by deliberately corrupting one pack: `restic check` catches it immediately („ciphertext verification failed", „Fatal: repository contains errors"), and the product only meets the damage when a customer is already trying to recover. | **OPEN — MEDIUM** | — | A periodic `restic check` (structure) with an occasional `--read-data`, reported like any other backup verdict. Note PBS already has verify jobs; this is the tier that does not. | CC | -| **R-360** | **The verification-copy delete gates on the concurrency flag, which the verification restore never holds — so the copy is deletable for the whole restore, and the handler's own comment claims the opposite.** `offboxVerifyCopyDeleteHandler` (`web/offbox_handlers.go:502`) reads `backupMgr.IsRunning()`; `RestoreOffboxScratch` (`offbox_restore.go:211`) **never calls `acquireRunning`**. R-351b moved all seven restore handlers onto `restoreOpBlocked()` (both flags) and left this one behind. Demonstrated 2026-08-21 22:35 with the flags read immediately before and after: `display=True offbox-restore kimai / concurrency=False` on both sides, and the delete succeeded. The guard does not compare app names, so the same call naming the restoring app removes the directory the restore is writing into. **Observed in two consecutive reports and filed neither time; filed now.** **CLOSED 2026-08-30, controller v0.226.0.** `restoreOpBlocked()` replaces `IsRunning()`, and **the doc comment that falsely claimed the handler already did this is corrected in place rather than deleted** — that sentence is why nobody looked. No app-name comparison was added: refusing during ANY restore is strictly stronger than refusing only for the restoring app, and it is the rule the other five handlers follow; uniformity is worth more than precision when the defect was one handler being different. **PROVEN LIVE on demo-hp 2026-08-30** in the exact state that produced the bug — a `mode=unit` restore in flight, display flag true and concurrency flag false — the delete was refused (`[WARN] verification-copy delete refused for kimai: a backup/restore op is running`) and **the planted canary file survived**, which is the consequence the unit test asserts too. Evidence: `documentation/audits/evidence-r353-r360-live-2026-08-30/live-validation.txt`. | **CLOSED — controller v0.226.0 (2026-08-30)** | — | `restoreOpBlocked()`, and a test that asserts the CONSEQUENCE — the copy survives a delete attempt mid-restore. | CC | | **R-362** | **A data drive detached mid-restore is reported as „permission denied".** Observed 2026-08-21 23:15: the guest-visible bind was unmounted 4 s into a scratch restore; the restore correctly failed and wrote nothing to the wrong place, but said „A visszaállítás sikertelen: restore dir: mkdir /mnt/felhom-drives/hdd_1/backups: permission denied". The controller has a drive-state concept (`IsDisconnected`, used by both backup legs) and the restore path never consults it. **A correct refusal that misdescribes why sends the reader at a permissions problem that does not exist.** Creditable in the same test: the agent re-bound the drive 5 s later, unaided. | **OPEN — MEDIUM** | — | Consult drive state when a restore path operation fails on ENOENT/EACCES and name the drive. | CC | | **R-363** | **The fill watcher runs once a day, so a filesystem that fills at 03:31 goes unannounced for ~24 h while the backup is already refusing apps.** `sched.Daily("fill-watch", "03:30", …)` (`cmd/controller/main.go:1092`) plus one startup check. Proven 2026-08-21 23:17: the 69 GB filesystem carrying the Docker data-root, the system namespace and ALL 40-class app data was filled to 99% / 1.2 GiB free; the backup reserve refused `kimai` per app and the hub received `recovery_unit_capture_failed` (error) naming the filesystem, **and the fill watcher said nothing at all**. The package comment says it "warns the CUSTOMER that a filesystem is filling, BEFORE anything fails"; at a daily cadence it frequently cannot. | **OPEN — MEDIUM** | — | The reserve already computes the same numbers every run. Let the watcher share that reading rather than owning a separate daily one. | CC | | **R-364** | **Accented-text search is an instrument that silently transforms its input, and discipline alone has failed at least three times.** (1) 2026-07-20, `ssh → pct exec → bash -c`, nearly a wrong "banner cleared" claim (`felhom-controller/.claude/rules/ui-hungarian.md:19-22`). (2) 2026-08-13, `kubectl exec … sh -c grep` returned **0 for three strings that were present**, one step from a wrongly-reported failed hub deploy. (3) 2026-08-21, `tar -tf` rendered `őszibarack.md` as `\305\221szibarack.md`; recording the fixture's name bytes from that listing would have been wrong. **NOTE: that is two inside two weeks plus the founding case a month earlier — a third inside the two-week window is not on record.** | **OPEN — LOW** | — | **PROPOSED, NOT BUILT:** a helper that refuses to report a zero for any pattern containing a byte ≥ 0x80 unless a negative control also returns zero AND an ASCII anchor known to be present returns non-zero. Three probes, one helper, no judgement at the call site — because judgement is what failed. | CC |