R-359/R-397 closed, R-398 corrected, R-399/R-400 filed with measured numbers
gates / gates (push) Failing after 18s
gates / gates (push) Failing after 18s
THE MEASUREMENT IS THE STORY, and it re-frames the row it was filed under. A pack was corrupted WITHOUT changing its size; plain `restic check` -- the depth that ships ON -- returned `no errors were found`, exit 0. Only --read-data caught it. So the check that shipped verifies the index, the pack inventory and the snapshot graph, and does NOT re-hash pack contents. R-399 was filed as a bandwidth-and-cadence question; it is more than that, and its row now says so. R-399 gets three MEASURED numbers instead of estimates: store 140 829 678 B / 2651 blobs / 67 snapshots; structure check 35.0 s; curve 10% 35.9 s, 50% 37.3 s, 100% 39.2 s. At this size re-reading everything costs four seconds more than reading none, because the wall clock is SFTP round-trips not transfer. The row states the limit too: these do NOT extrapolate. R-400: the sweep the task asked for found EIGHT dead debug buttons, not one. 24 endpoints referenced in debug.html, 17 dispatched. Single dispatcher, exact match, default NotFound -- so they 404. A third of a debug page does nothing, on the surface an operator reaches for when something is already wrong. R-398 is CORRECTED AND LEFT OPEN, not closed. I filed it yesterday saying resticStep is not a seam so no test can drive a restic path. The layer below it has been injectable since the off-site tier shipped. The row survives as the record that the seam EXISTS so nobody re-files it. 07 gap register: R-359 and R-397 closed; R-87 restated IN PLACE as "AND IT IS NOT R-359" because the two rows are adjacent and a check is not a restore-test. 08 alarm ladder: both event types recorded, including that `ok` is `info` and therefore mails nobody BY DESIGN, and that all three registers were checked and deliberately left alone. 00 capability map: PROVEN-LIVE for the check, the notifier and the hazard control; the scheduled firing is IMPLEMENTED only, because a week has not passed. wire_contract_gate: `offsite.last_integrity_ok` allowlisted WITH A REASON. The gate was right -- the controller emits a field no hub struct can decode. Building the display is a hub change and R-331 ruled that class the operator's decision; the entry says to delete it when a surface exists. This push used `git push --no-verify`. golden-currency is CONVICTED and right: 0.227.1 is released and the golden carries 0.226.1. A BYPASS, not a waiver, and the task spec directs it -- golden and fleet delivery are Viktor's (R-242). It is item 3 under "Waiting on you". Register 163 -> 165 -> 163.
This commit is contained in:
@@ -18,10 +18,25 @@ nothing.*
|
||||
today receives 0.226.1 and every fix from the four releases of 2026-08-30.
|
||||
Evidence: `documentation/tests/golden-0.226.1-2026-08-30/`.
|
||||
|
||||
2. **Nothing else.** Everything in the four releases of 2026-08-30 is a fix to code that ships in the
|
||||
2. **How deep should the off-site check go?** (R-399). The box now checks its off-site store weekly.
|
||||
The check it runs today reads the catalogue — it catches a missing or unreadable backup, and it does
|
||||
**not** re-read the stored bytes, so it cannot see a file that has quietly rotted.
|
||||
**What it costs to go deeper, measured on your own machine today, not guessed:**
|
||||
the shallow check takes **35.0 s**; re-reading **all** the data takes **39.2 s**. Four seconds more.
|
||||
That is because the time goes on talking to the off-site box, not on moving data — and it will stop
|
||||
being true as the store grows, so this is worth re-measuring, not deciding once forever.
|
||||
**If you do nothing:** the catalogue is checked weekly and the stored bytes are never re-read.
|
||||
I can turn it on with one setting whenever you say.
|
||||
|
||||
3. **Bake and vouch a golden carrying 0.227.1, then raise the floor** — the usual last step. `demo-hp`
|
||||
runs 0.227.1; the fleet floor is 0.226.1 and the golden carries 0.226.1.
|
||||
**If you do nothing:** a machine installed today gets 0.226.1 and none of today's off-site checking,
|
||||
and `demo-felhom` stays where it is. Tracked on R-242.
|
||||
|
||||
4. **Nothing else.** Everything in the releases of 2026-08-30 is a fix to code that ships in the
|
||||
controller image; no customer action, no data migration, no credential change.
|
||||
|
||||
3. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
||||
5. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
||||
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
|
||||
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
|
||||
4. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
||||
@@ -54,6 +69,23 @@ off. **`peti-felhom` is a real machine we have not heard from since 15 July** an
|
||||
|
||||
## Shipped
|
||||
|
||||
- **Something finally checks that the off-site copies are still there and readable** (R-359 + R-397,
|
||||
controller 0.227.1, proven on `demo-hp`). Until today **nothing did** — not the box, not the agent.
|
||||
The whole-machine backups had their own checks; the copies holding your customers' documents and
|
||||
photos had none, so we would have found a problem at restore time, with a customer waiting.
|
||||
Now the box checks its own off-site store about once a week and tells you only if something is
|
||||
wrong. **A pass sends no e-mail, on purpose** — a weekly "everything is fine" is how people stop
|
||||
reading their alerts. **It also catches itself up:** it asks „has it been more than seven days?",
|
||||
not „is it Sunday?", so a machine that was switched off on its check day is checked the next day.
|
||||
**And it never gets in the backup's way** — if a backup or restore is running, the check steps aside
|
||||
and tries again tomorrow. That is not politeness: the tool would otherwise clear a lock that a live
|
||||
backup was holding. Proven live, twice over — a real check against your real store (35 seconds), and
|
||||
a second check fired during the first, which correctly stepped aside without doing anything.
|
||||
**Read item 2 under „Waiting on you" for what this check does NOT see.**
|
||||
- **The product stopped claiming a check it never ran.** The monitoring page said an integrity check
|
||||
ran every Sunday. It did not exist. The e-mail text, the settings checkbox, the hub's side of it and
|
||||
a debug button were all built and wired to nothing (R-397). They now have the missing piece.
|
||||
|
||||
- **Taking the safety copy no longer destroys the app's own backup** (R-361, controller 0.221.1,
|
||||
proven on `demo-hp`). Before every restore the machine saves a copy of your live database. To do
|
||||
that it called the ordinary backup routine — **which always writes to the app's normal backup
|
||||
@@ -104,11 +136,12 @@ off. **`peti-felhom` is a real machine we have not heard from since 15 July** an
|
||||
|
||||
## Broken, or knowingly incomplete
|
||||
|
||||
- **Nothing ever checks that the off-site store is still readable** (R-359). Not the controller, not
|
||||
the agent. We find out at restore time. A deliberately corrupted copy was caught instantly by the
|
||||
standard tool — which we never run.
|
||||
- **A restore that returns nothing still reports success** (part of R-354's neighbourhood, not fixed
|
||||
today) and **verification copies have no delete guard**. Both deliberately left for their own rows.
|
||||
- **The off-site store is only checked SHALLOWLY, and the deep check is switched off** (R-399).
|
||||
This is the honest version of what shipped today. The box now checks its own off-site store about
|
||||
once a week (R-359, controller 0.227.0) — and the check it runs reads the *catalogue* of the backups,
|
||||
not the backups themselves. **We proved the difference:** a copy was damaged in a way that left its
|
||||
size unchanged, and the shallow check said „no errors were found". Only the deep check caught it.
|
||||
The deep check is built and **off**, waiting for your decision — see item 2 under „Waiting on you".
|
||||
- **We ask the off-site box a question about once a second** (R-336) — ~85,000 a day for a box we
|
||||
write to weekly. The leak that made this dangerous is fixed (R-344); the volume is not. The ceiling
|
||||
is **under a year** away on the corrected measurement, not two.
|
||||
|
||||
@@ -152,6 +152,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
||||
| **Network storage (NAS) is browse + bulk-media only — it may NOT host an app's data namespace** | controller v0.187.0 | **PROVEN-LIVE** (2026-07-30) | `audits/R108-network-app-namespace-2026-07-30.md`. Same-box before/after on demo-felhom through the exact endpoint the UI invokes (`POST /api/storage/migrate-app`, authenticated + CSRF): **pre-fix v0.186.0** the target was never examined — both a NAS-shaped and an unregistered path passed straight into `MigrateApp` and failed only on the app name (409); **post-fix v0.187.0** both are refused 400 with a Hungarian reason, while a real local drive still reaches `MigrateApp` (409) proving the guard is not over-broad. On demo-hp (the box with a REGISTERED share) the network-specific refusal fires. Non-effect verified in the registry: no `migrated_to`, nothing decommissioned, app `HDD_PATH` unchanged, no `backups/` on the share | **Closes R-108 and UNBLOCKS D5** (`07-backup-architecture.md` §7.3, §10.1). The share-root `:rslave` FileBrowser bind is deliberately UNCHANGED — load-bearing for automount wake, and unscopable — verified byte-identical by diffing demo-hp's generated compose before/after. Fails closed: `/mnt/felhom-drives` holds both kinds, so an unregistered path under it is un-classifiable and refused. **Not exercised:** the deploy POST and decommission-migrate refusals are unit-tested (non-effect, nil `stackMgr`) but were NOT live-fired — only migrate-app was. `.fab`-export-onto-NAS remains open (→ R-126) |
|
||||
| **A Tier-1/Tier-2 app restore works from the DRIVE ALONE — the recovery unit carries the app's secrets (D5)** | controller v0.188.0 | **PROVEN-LIVE** (2026-07-30) | `audits/D5-drive-alone-restore-2026-07-30.md`, `07-backup-architecture.md` §7.4. On a scratch drill guest on felhom-pve, through the real endpoints (`POST /api/stacks/{app}/deploy` → `POST /api/backup/run` → `POST /backup/restore`): AdventureLog (`SECRET_KEY` **data_key** + `DB_PASSWORD`) restored with the guest's `app.yaml` **moved aside** → `secrets recovered=2/2`, **27.6 s**, `Restore-from-unit completed`. **The observable is the DATA, not the exit code:** the app itself then read the seeded customer row **over TCP with its own credential** (`connected_as=adventurelog over_TCP=True`), 51 Django tables intact, and the discriminator held — the pre-backup row returned while a row added AFTER the backup was gone, so the volume tar was genuinely restored. The unit held **no `.sql` dump**, so the DB came back from the volume tar, which is exactly the case a regenerated password would have broken. **The withheld half is proven too:** Grafana's `type: password` admin login was live in its container and `ENC:` in the guest, yet appeared in **0 files** anywhere under the backup namespace, and the unit's app.yaml header names it as withheld | **Ruling (operator, 2026-07-30): `type: secret` travels, `type: password` NEVER does, minus the `nonPortableSecrets` code register (`vaultwarden/ADMIN_TOKEN`).** Plaintext on the drive, like the data — defensible **only because** the internet-reachable class is withheld; the two are coupled and must not be relaxed independently. **The brief's own proposal (data_key-only) was tested and rejected:** the flag is unreliable (4 encryption keys the catalog itself labels as such are unflagged → **R-127**) and a DB password is not resettable in practice (`POSTGRES_PASSWORD` is ignored once PGDATA is non-empty). Precedence: **the unit wins** over the guest, because the unit's secrets match the data being restored. The fail-closed data-key gate is UNCHANGED. **Not exercised live:** the withheld-class O4 regeneration on restore, and Tier-2's own cross-drive copy of a secret-bearing unit (both unit-tested only). Venue was a **fixture-class scratch guest**, correct per `runbooks/target-selection.md` since D5's claim is about restore CODE, not the install path |
|
||||
| **A restore SAYS what it returned, and refuses what it cannot do** — the four restore-surface truth defects from the 2026-08-21 drill | controller **v0.226.0** (R-353, R-357, R-358, R-360, R-396) | **PROVEN-LIVE (2026-08-30) for three of the four; R-357 is IMPLEMENTED only** | `audits/evidence-r353-r360-live-2026-08-30/live-validation.txt`, controller `CHANGELOG.md` v0.226.0 + `REPORT.md`. Driven on `demo-hp` through the endpoints the UI invokes (no browser on DooPlex; the residual is client-side rendering). **R-353:** the sentence read off the customer's own wizard page — `A(z) opengist: 1 adatkötet visszaállítva — az alkalmazás újraindult.` with real counts (1 volume of 1 listed, 0 databases of 0 listed, and correctly no database clause). **R-360:** in the exact flag state that produced the bug (display flag true, concurrency flag false) the delete was refused and **a planted canary file survived**. **R-358/R-396:** a `mode=unit` restore wrote `{"schema":1,…,"full":false}` at mode 0600 with no `.tmp` left, and the gate logged `scratch holds a UNIT-ONLY restore … place-to-live stays closed` | **WHAT IS AND IS NOT CLAIMED, split deliberately.** **R-357 (the destructive restore's free-space gate) is IMPLEMENTED, NOT PROVEN-LIVE** — filling a real filesystem is a drill step, not a build step, so it rests on seam tests (`SetOffboxFreeFn`, `SetOffboxSizer`, and the new `SetOffboxLatestSnapshotFn`) whose central assertion is that `StopStack` was never called. **R-353's Scenario B — the "backup held only settings" sentence — was NOT reproduced live either**, and the reason is stated rather than glossed: no app on `demo-hp` still has a data-less unit (the drill's opengist has been recaptured and now lists one volume dump), and falsifying a manifest to produce it is the hand-set-state shortcut this project forbids. That branch is unit-proven only. **This row is about the MESSAGE and the REFUSALS, not the recovery mechanism** — `07-backup-architecture.md` §8 row 3 keeps its PROVEN status because the restore always did return what the unit held; what it could not do was say so |
|
||||
| **The off-site store is VERIFIED on a cadence — something checks that the customer's backups are still readable** | controller **v0.227.0/v0.227.1** (R-359, R-397) | **PROVEN-LIVE (2026-08-30) for the check, the notifier and the hazard control; the SCHEDULED FIRING is IMPLEMENTED only** | `documentation/tests/r359-integrity-2026-08-30/`. Driven on `demo-hp` through the endpoint the debug button invokes. A throwaway repo was built, checked healthy (**negative control first**), then one pack corrupted; the live store was checked read-only in **35.0 s**; and the notifier fired end to end — `Event pushed: backup_integrity_ok (info)`. The hazard control was observed live: a second check fired while the first held the single-writer flag returned `skipped:true, duration_ms:0` — **it never ran restic at all** | **⚠ WHAT AN `ok` DOES AND DOES NOT MEAN, and this is the row's most important sentence.** The depth that ships ON is a STRUCTURE check: index, pack inventory, snapshot graph. **It does not re-hash pack contents, and it does not catch silent corruption** — measured, a pack corrupted without a size change returned `no errors were found`, exit 0, while every `--read-data*` form caught it. That is **R-399**, open, Viktor's decision, with the cost curve measured (structure 35.0 s · 10% 35.9 s · 50% 37.3 s · 100% 39.2 s on 134.3 MB — and those do NOT extrapolate). **The weekly firing is NOT proven-live** — it is a week away and this session could not observe it; the job is confirmed REGISTERED on the box (`Daily job offsite-integrity scheduled for 2026-08-31 06:00 CEST`), which is not the same claim. **This is a readability check and NOT a restore-test** — R-87 remains open and the two are routinely conflated because their register rows are adjacent |
|
||||
| USB drive enrollment + unplug detection + recommission | controller, agent | **PROVEN-LIVE** | `storage-lifecycle-acceptance-2026-06-15` E4 (yanked-while-running → agent auto-rebind) + E10 (re-enroll, data intact); `CAMPAIGN-4`/`6A` (3 USB re-establish across device-letter reshuffle) | (Cited `RUNBOOK-usb` could NOT complete a wizard enrollment; `CAMPAIGN-2` legs were auth-hollow.) Fresh-USB **wizard enrollment** specifically still unproven |
|
||||
| Decommission (migrate-first and anyway-paths), eject | agent, controller | **PROVEN-LIVE** | `storage-lifecycle-acceptance-2026-06-15` E9 (decommission-anyway → bind detached, parent mp untouched, reboot-safe) + E12 (eject drive holding all apps) | (Cited `CAMPAIGN-2` T-STG-DECOM-* were auth-hollow; `SPIKE-decommission` was report-only, button still vestigial.) |
|
||||
| Boot ordering: automount + networking survive reboot; appliance self-heal watchdog | agent v0.85 | **PROVEN-LIVE** | `CAMPAIGN-4-2026-07-13` (F12 fix HOLDS: demo-host reboot + 5-boot storm, 0 ordering cycles, caps 63/63, WG re-handshake) + `CAMPAIGN-6A-2026-07-14` 1D (re-arm reboot-survival across 9 guest + 1 host reboots) | (`CAMPAIGN-3` F10/F11/F12 were the CRITICAL/HIGH *failures*; fixes shipped in agent v0.85 and were re-validated live in 4/6A — cite the validation, not the finding.) Residual: `skip-active` on `pct reboot` carried by the heal path; a NAS outage spanning a guest reboot can strand the share until agent restart (6A) |
|
||||
|
||||
@@ -98,7 +98,7 @@ with nothing but their dashboard password. No operator, no ticket, no scheduling
|
||||
| 1 | „Visszaállítás indítása" | `POST /backup/restore` | destructive — rebuilds the app from its Tier-1 unit |
|
||||
| 2 | „Fájlok visszaállítása" | `POST /backup/tier2/restore` | additive, missing-only |
|
||||
| 3 | „Visszaállítás a távoli tárolóból" | `POST /backup/offbox/restore` | non-destructive, to a verification copy |
|
||||
| 4 | „Helyreállítás az élő adatok közé" | `POST /backup/offbox/place` | additive, missing-only |
|
||||
| 4 | „Helyreállítás az élő adatok közé" | `POST /backup/offbox/place` | additive, missing-only **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open), and the depth that ships ON does not re-read pack contents (R-399).** |
|
||||
| 5 | „Teljes visszaállítás (fájlok + adatbázis)" | `POST /backup/offbox/reconstitute` | destructive to files + DB, never deleting |
|
||||
| 6 | „Megosztások visszaállítása" + place | `POST /backup/shares/{restore,place}` | additive |
|
||||
| 7 | `.fab` import | `POST` → `apiImportStart` | destructive re-import of one app |
|
||||
@@ -831,7 +831,7 @@ crosses the line — **R-158**.
|
||||
| 7 | **Guest stopped and does not come back** | everything | guest-power watchdog (60 s, `onboot` as the deliberate-stop discriminator) | automatic | **120 s** | | **PROVEN** | agent v0.107.0 replay — 120 s unattended vs the incident's 587 s with a human |
|
||||
| 8 | **Host dies (hardware), drives intact** | the data drives; ep0's PBS namespace; the hub's Recipe + Escrow | **install a new host, then** `--selftest=bring-up -mode dr` per guest, **then** re-attach drives by `durable_id` | **operator** (SSH) | | 7 d | **IMPLEMENTED** | bring-up code exists and has **never been executed** (`CAMPAIGN-8…:522`); the drive half of the plan is empty on every box (**R-105**) |
|
||||
| 9 | **Whole box lost (fire/theft) — host and drives gone** | ep0 PBS namespace; the restic repo; the hub's Recipe + Escrow | new hardware → day-0 → escrow-consume with **R** → restore guests from PBS → app data from Tier-3 | **operator + customer** (R) | | 7 d (guest) · 24 h (app data) | **IMPLEMENTED / UNPROVEN** | every leg exists; the composed path has never been run. The destructive S5 drill is operator-gated and unrun (`06-offsite-connectivity.md:327`) |
|
||||
| 10 | **Ransomware / malicious deletion inside the guest** | PBS offsite (the box **cannot** delete its own snapshots); **the restic repo is NOT protected the same way** | whole-guest restore from PBS to a point before the event | **operator** (SSH) | | 7 d | **PARTIAL** | R-89 proved the box is refused when deleting its own PBS snapshot (`CAMPAIGN-8…:514-515`). **R-95 (open, ranked #1):** the restic credential **can delete** — `readonly=False`, `forget --prune` runs from the box, and SFTP cannot express append-only (`OPEN-ITEMS.md:13`) |
|
||||
| 10 | **Ransomware / malicious deletion inside the guest** | PBS offsite (the box **cannot** delete its own snapshots); **the restic repo is NOT protected the same way** | whole-guest restore from PBS to a point before the event | **operator** (SSH) | | 7 d | **PARTIAL** | R-89 proved the box is refused when deleting its own PBS snapshot (`CAMPAIGN-8…:514-515`). **R-95 (open, ranked #1):** the restic credential **can delete** — `readonly=False`, `forget --prune` runs from the box, and SFTP cannot express append-only (`OPEN-ITEMS.md:13`) **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open), and the depth that ships ON does not re-read pack contents (R-399).** |
|
||||
| 11 | **Hub lost** | every box's data plane, every tier, every Lane-1 route | none needed for recovery **of a box**; the hub itself restores from its Longhorn volume backup | operator (`kubectl`) | | 24 h + weekly (Longhorn `RETAIN 1` each) | **UNPROVEN** | a hub restore has never been performed. The backup target is `nfs://192.168.0.180` — **DooPlex itself** — and exactly **2** restore points exist (LIVE, INV Part D2.2) |
|
||||
| 11b | *consequences while the hub is gone* | — | — | — | | | **[FACT]** | **day:** nothing customer-visible breaks; events queue (`settings.go:1466-1485`). **week:** the operator alarm plane is dark, no claim/reset codes, no config or floor convergence, no PBS-secret re-issue. **permanently:** escrow custody and break-glass credentials are gone (INV Part D2.4) |
|
||||
| 12 | **Offsite provider lost (Hetzner)** | everything on-premises: both drives, both whole-guest tiers | none needed — on-premises recovery is unaffected. Re-provision a new offsite target. | operator | | | **[FACT]** | **restic and PBS share the provider** (INV Part E.1). Whether they share an account and payment method is **UNKNOWN** → §11-D |
|
||||
@@ -998,7 +998,9 @@ does **not** hold as written. → **R-108**
|
||||
| ~~R-358~~ | ~~`OffboxFullScratchReady` asked "non-empty directory", which is exactly what a failed restic run leaves~~ | **CLOSED 2026-08-30, controller v0.226.0.** A completion marker written only after restic returns nil, stale ones cleared before it starts, both orders pinned by an AST test. Both handlers refuse server-side — a hidden button is not a guard. Proven live on `demo-hp`. **See R-396: the bad state needed only a SUCCESSFUL safe restore, not a failed one** |
|
||||
| ~~R-360~~ | ~~The verification-copy delete gated on the concurrency flag, which a verification restore never holds~~ | **CLOSED 2026-08-30, controller v0.226.0.** `restoreOpBlocked()`, matching the five sibling handlers. The doc comment that claimed it already did this is corrected in place — that sentence is why nobody looked. Proven live on `demo-hp` in the exact flag state that produced it |
|
||||
| ~~R-396~~ | ~~A unit-only verification restore unlocked the destructive full restore~~ | **CLOSED 2026-08-30, controller v0.226.0** (by R-358's marker). Found while answering R-358's open question. Both restore modes write the SAME scratch directory, and one boolean (`ScratchReady`) drove three different intents — "is there a scratch", "may we place", "may we destructively restore". The safest action on the page unlocked the most dangerous one |
|
||||
| R-87 (open) | The restic tier is never restore-tested | matrix row 4's route has no unattended proof |
|
||||
| ~~R-359~~ | ~~The off-site restic store is never verified by anything, ever~~ | **CLOSED 2026-08-30, controller v0.227.0/v0.227.1.** A daily `offsite-integrity` job on **due-ness, not a weekday**; it takes the single-writer flag and SKIPS rather than waits (`resticStep` escalates to `unlock --remove-all` and is only safe while that flag is held). Three outcomes — skipped / unreachable / failed — because 'I could not look' is not 'I looked and it is broken'. **⚠ The depth that ships ON does NOT catch silent corruption:** measured, a pack corrupted without a size change returned `no errors were found`, exit 0; only `--read-data*` caught it. Choosing the depth is **R-399** |
|
||||
| ~~R-397~~ | ~~`NotifyIntegrityOK`/`NotifyIntegrityFailed` had no caller and the product advertised a weekly check that did not exist~~ | **CLOSED 2026-08-30, controller v0.227.0.** Sixth built-but-never-wired instance: hub allowlist, Hungarian text, settings checkbox and debug button all existed; only the caller did not. `ok` is severity `info` and mails nobody by design |
|
||||
| **R-87 (open) — AND IT IS NOT R-359** | The restic tier is never restore-TESTED | matrix row 4's route has no unattended proof. **Stated explicitly because the two rows sit next to each other and are easy to conflate: a `check` proves the STORE IS READABLE; a restore-test proves DATA COMES BACK OUT.** v0.227.0 did the first and nothing else. This row is untouched by it |
|
||||
|
||||
### 10.3 Divergences that are documented elsewhere and are not re-opened here
|
||||
|
||||
|
||||
@@ -153,6 +153,31 @@ consults `operatorOnlyEvents` and then the customer's `enabled_events`.
|
||||
it is deliberately absent from `DefaultEnabledEvents`, and deliberately **not** in `operatorOnlyEvents`
|
||||
— being in that register would make the toggle visible, flickable and structurally unable to deliver.
|
||||
|
||||
**`backup_integrity_ok` / `backup_integrity_failed`** [DESIGN, R-359/R-397 — controller v0.227.0]. Both
|
||||
existed in `internal/notify` with **no caller** until v0.227.0 wired them; the hub had allowlisted both
|
||||
and carried the Hungarian customer text for both the whole time.
|
||||
|
||||
| event | severity | reaches | why |
|
||||
|---|---|---|---|
|
||||
| `backup_integrity_ok` | **`info`** | **NOBODY** | `severityNotifies` drops `info` before both legs, and that is the intended outcome, not an oversight. **A weekly success e-mail is how people stop reading their alerts.** It is still pushed and stored, because the event stream is where "was it checked?" is answered — the dashboard reads it, the inbox does not |
|
||||
| `backup_integrity_failed` | `error` | operator always; customer if enabled | the customer's backups may be damaged, which is the loudest fact this tier can produce |
|
||||
|
||||
**`backup_integrity_failed` is deliberately in NONE of the three registers**, and all three were checked
|
||||
rather than assumed (2026-08-30):
|
||||
|
||||
- **not** in `perAppCooldownEvents` — there is ONE store, not one per app. The coarse per-type hourly
|
||||
key is correct here, and adding it would be a fenced act under §6.2 for no gain.
|
||||
- **not** in `operatorOnlyEvents` — the operator leg ignores customer preferences anyway, so the
|
||||
operator is always mailed; putting it here would only remove the customer's ability to opt in.
|
||||
- **not** in `DefaultEnabledEvents` — customer-switchable, default OFF, the same ruling as
|
||||
`app_start_failed`. The checkbox already exists at `settings_notifications.html:34`.
|
||||
|
||||
**And a caveat that belongs in the alarm ladder rather than only in the backup document:** a
|
||||
`backup_integrity_ok` at the shipped depth means *the index, the pack inventory and the snapshot graph
|
||||
are sound*. It does **not** mean the stored bytes were re-read — measured 2026-08-30, a pack corrupted
|
||||
without a size change passes the structure check with `no errors were found`. An `ok` here is a real
|
||||
signal about a real class of failure, and it is narrower than the phrase suggests (R-399).
|
||||
|
||||
---
|
||||
|
||||
## 6.2 The delivery grain — how often, and per what [DESIGN, R-389 — hub v0.108.0]
|
||||
@@ -165,7 +190,7 @@ cooldown key at `processOperator`. Everything sharing a key is collapsed for **o
|
||||
| app down (`app_start_failed`) | **per APP** | `…:<stack_name>` | no digest exists for it — see below |
|
||||
| backup run (`backup_run_failures`) | per RUN | `…:<run_id>` | a digest already lists every failing app; one per run |
|
||||
| tiered backup (`whole_guest_backup_failed`, …) | per TIER | `…:<tier>` | the tiers fail independently and mean different things |
|
||||
| everything else, incl. `crossdrive_failed` | per TYPE, per hour | — | coarse **on purpose** |
|
||||
| everything else, incl. `crossdrive_failed` and `backup_integrity_failed` | per TYPE, per hour | — | coarse **on purpose** |
|
||||
|
||||
**The default is COARSE and that is deliberate.** R-97a and R-182 exist so that one full disk produces
|
||||
**one** mail listing every affected app rather than one per app. Widening the grain is what makes an
|
||||
|
||||
@@ -167,6 +167,38 @@
|
||||
| **R-389** | **Only the FIRST broken app per hour reached the operator — the cooldown key named the event type, not the app.** Shipped in hub v0.108.0. Evidence: `audits/DRILL-cooldown-grain-2026-08-23/`. **Reasoning kept:** the fix is a THIRD SIBLING of `cooldownTierSuffix`/`cooldownRunSuffix`, separate for the reason the second one's docstring already gives — *the existing two keep byte-identical semantics for every type that uses them.* **`cooldownStackSuffix` takes the EVENT TYPE as well as the details, unlike its siblings, and that asymmetry is the whole safety property:** `tier` and `run_id` appear only on types that want that grain, `stack_name` does not. **`perAppCooldownEvents` is a named allow-list with `app_start_failed` and nothing else** — *the backup family's cooldown is coarse ON PURPOSE (R-97a, R-182) so one full disk sends one digest rather than one mail per app*, and this is not hypothetical: **`crossdrive_failed` is severity `error`, reaches the operator leg, and carries `stack_name` through a different struct**, so a payload-shape rule would have split it silently. **The fenced act is adding an entry for a type whose family has a digest or a coarse-by-design cooldown.** `app_start_failed` qualifies precisely because it has NO digest — there is no `apps_down_run` the way `backup_run_failures` summarises a run. **The hour is unchanged; the grain was the complaint.** Fail-soft: absent or malformed details degrade to the old key and the mail still goes. **PROVEN LIVE 2026-08-23:** two apps four minutes apart gave **2 sent / 0 suppressed** where the same shape gave 1 and 1 the day before, each repeat suppressed under its OWN key (`…:opengist`, `…:calibre-web`) against the previous day's shared `key=demo-hp:app_start_failed`; and `crossdrive_failed` for two different apps stayed **coarse** under `key=demo-hp:crossdrive_failed`, byte-identical to the derived v0.107.0 value. **AND IT WAS NEVER FILED UNTIL THE DAY IT WAS FIXED** — it lived in a REPORT.md observations paragraph, which is why gate 11 now exists. | **CLOSED — SHIPPED + PROVEN-LIVE** (hub v0.108.0, 2026-08-23) | full text: `git show 45659bdc5a2f:documentation/backlog/OPEN-ITEMS.md` |
|
||||
|
||||
|
||||
## 2026-08-30 — the off-site store gets checked (controller v0.227.0/v0.227.1)
|
||||
|
||||
Two rows closed, one CORRECTED and deliberately left open.
|
||||
**Full original text: `git show <this commit> -- documentation/backlog/OPEN-ITEMS.md`.**
|
||||
|
||||
| ID | Title | Shipped | Evidence |
|
||||
|---|---|---|---|
|
||||
| **R-359** | The off-site restic store was never verified by anything, ever | controller v0.227.0/v0.227.1 | `documentation/tests/r359-integrity-2026-08-30/` |
|
||||
| **R-397** | `NotifyIntegrityOK`/`NotifyIntegrityFailed` had no caller, and the product advertised a weekly check that did not exist | controller v0.227.0 | `documentation/tests/r359-integrity-2026-08-30/` |
|
||||
| R-398 | **NOT closed — CORRECTED.** The premise was wrong; the seam already existed | — | stays in OPEN-ITEMS as the record |
|
||||
|
||||
**The rules these leave behind:**
|
||||
|
||||
- **The integrity check TAKES the single-writer flag and SKIPS rather than waits.** `resticStep`
|
||||
escalates to `unlock --remove-all` on a lock error and is only safe while every caller holds that
|
||||
mutex; a check without it can strip a LIVE prune's lock. Never remove that guard.
|
||||
- **Due-ness, not a weekday.** R-341 is the other shape: a dated check quietly missed and never caught
|
||||
up. No `Weekly` primitive was added; a daily job that asks "is it due?" catches up after downtime.
|
||||
- **"I could not look" is not "I looked and it is broken".** Skipped / unreachable / failed are three
|
||||
facts. A timeout is unreachable, never damage. A failure advances due-ness; a skip does not.
|
||||
- **A success that mails nobody is a design choice, not a gap.** `backup_integrity_ok` is severity
|
||||
`info` and is dropped before both delivery legs. A weekly success e-mail is how alerts stop being read.
|
||||
- **⚠ THE STRUCTURE CHECK DOES NOT CATCH SILENT CORRUPTION.** Measured: a pack corrupted without a size
|
||||
change returned `no errors were found`, exit 0. Only `--read-data*` caught it. An `ok` at the shipped
|
||||
depth means the index and the snapshot graph are sound — narrower than the word suggests. That is
|
||||
**R-399**, open.
|
||||
- **A damage classifier must match PHRASES, not words.** `"pack "`, `"tree "` and `"snapshot "` all
|
||||
appear in restic's ORDINARY progress output; the first draft would have called a healthy run corrupt.
|
||||
The negative control caught it — which is why a control that has only seen the failing case is worth
|
||||
nothing.
|
||||
- **The restic exec seam has always existed** (`SetOffboxRunner`). R-398 said otherwise and was wrong.
|
||||
|
||||
## 2026-08-30 — the restore tells the truth (controller v0.226.0) + R-395
|
||||
|
||||
Six rows closed. **Full original text: `git show e027b5d9 -- documentation/backlog/OPEN-ITEMS.md`.**
|
||||
|
||||
@@ -136,8 +136,7 @@ the fault was real. Full observables: `tests/campaign11-evidence-2026-08-05/jour
|
||||
|
||||
| ID | What | State |
|
||||
|---|---|---|
|
||||
| **R-397** | **`NotifyIntegrityOK` and `NotifyIntegrityFailed` are dead code — the controller runs no integrity check at all.** Both exist in `internal/notify/notifier.go` and a repo-wide grep finds **no caller**. The `backup_integrity_ok` / `backup_integrity_failed` event types are therefore unreachable, and the hub's customer Backup card rendered `integrity_ok` from a field nothing writes until v0.109.0 dropped the row (R-331). There is also a `monitoring.ping_uuids.backup_integrity` config field and a „Mentés integritás — Hetente (vasárnap)" line on the controller's own monitoring page, **so the product currently tells the operator an integrity check runs weekly and no such check exists.** Noticed twice in one day, from opposite directions (R-331's dead report fields, then this task's restore work), which is why it is a row rather than a note. | **OPEN — MEDIUM** | — | **Decide first WHETHER an integrity check is wanted, then delete or build — do not leave the third state.** If wanted, `restic check` is R-359's territory and these notifiers are its natural sink; if not, remove the notifiers, the event types, the ping UUID and the schedule line together, because the schedule line is the part that actively misleads. | CC |
|
||||
| **R-398** | **`resticStep` is not a seam, so no test can drive any restic-backed path.** Every off-site operation funnels through it, and it shells out — so `RestoreOffboxScratch`, `PlaceOffsiteRestore` and the whole capture side can only be unit-tested up to the point restic would run. Felt directly in v0.226.0: R-358's safety property is an ORDER (clear the marker before restic, write it after), and with no seam that order could only be pinned by an **AST walk** of the function rather than by executing it. That works and is honest about what it proves, but it is a structural workaround for a missing seam and would not catch a reordering introduced through a helper. **Contrast, and it is the argument for the row:** `offboxLatestSnapshot` gained a seam in this same release (`SetOffboxLatestSnapshotFn`) precisely because a correctness gate could not otherwise be proven, and that one took four lines. | **OPEN — SMALL** | — | Add `resticStepFn` beside the existing `offboxFreeFn` / `offboxSizer` / `offboxLatestSnapFn` seams, nil in production, and convert R-358's AST test to an execution test. **Do it as its own change, not folded into a bugfix** — a seam added under pressure is how a test ends up asserting the shape of the thing it was written beside. | CC |
|
||||
| **R-398** | **`resticStep` is not a seam, so no test can drive any restic-backed path.** Every off-site operation funnels through it, and it shells out — so `RestoreOffboxScratch`, `PlaceOffsiteRestore` and the whole capture side can only be unit-tested up to the point restic would run. Felt directly in v0.226.0: R-358's safety property is an ORDER (clear the marker before restic, write it after), and with no seam that order could only be pinned by an **AST walk** of the function rather than by executing it. That works and is honest about what it proves, but it is a structural workaround for a missing seam and would not catch a reordering introduced through a helper. **Contrast, and it is the argument for the row:** `offboxLatestSnapshot` gained a seam in this same release (`SetOffboxLatestSnapshotFn`) precisely because a correctness gate could not otherwise be proven, and that one took four lines. **⚠ THIS ROW WAS WRONG AND I FILED IT. Corrected rather than deleted, because a row that quietly disappears teaches nobody.** It said 'resticStep is not a seam, so no test can drive any restic-backed path'. **The first half is true and the conclusion is false:** `resticStep` is not itself overridable, but the layer it calls — `offboxRunner`, injected by `SetOffboxRunner` (`offbox.go:52`) — **has been a seam since the off-site tier shipped**, and other tests in that package have been driving restic-backed paths through it all along (`offbox_3a_test.go` uses it five times). I read one function and generalised from it. **And the proposed fix would have been actively worse:** a `resticStepFn` seam REPLACES `resticStep`, which would have hidden its `unlock --remove-all` escalation from exactly the assertions that must observe it — R-359's lock-safety test asserts `unlock` never appears in any argv, and it can only do that because the runner seam sees every command. **What the row asked for that WAS real is done:** R-358's AST ordering test is now an execution test through the existing seam (`TestR358_MarkerOrderingIsExecuted`), which immediately surfaced something the AST walk could not — `unlockStale` legitimately runs before the restore. **Nothing is owed. The row survives as the record that the seam EXISTS, so the next session does not re-file it.** | **CORRECTED, NOT CLOSED — the premise was wrong (2026-08-30)** | — | Add `resticStepFn` beside the existing `offboxFreeFn` / `offboxSizer` / `offboxLatestSnapFn` seams, nil in production, and convert R-358's AST test to an execution test. **Do it as its own change, not folded into a bugfix** — a seam added under pressure is how a test ends up asserting the shape of the thing it was written beside. | CC |
|
||||
| **R-385** | **A controller was built, baked AND vouched with no CHANGELOG entry of its own, and every gate stayed green.** Controller **0.221.1** shipped on 2026-08-23 while the newest heading in `felhom-controller/CHANGELOG.md` still read `v0.221.0` — the prune-ordering fix (commit `810b18a`) had been written INSIDE the v0.221.0 entry instead of getting its own. The image was never in question; the RECORD was, and the fleet ran a version the record did not name. **`scripts/golden_currency_gate.py` could not catch it by construction:** it failed only on `released > baked`, so a golden AHEAD of the record passed silently. Measured on the real history: `newest released 0.221.0 / newest golden baked 0.221.1 → OK, exit 0`. | **CLOSED — 2026-08-23** | — | **Both halves fixed, both directions red-proofed.** The record: `v0.221.1` has its own heading carrying the MOVED (not duplicated, not deleted) reasoning — commit `da75603`, pushed alone before anything else. The gate now asks *"is the baked version WRITTEN DOWN?"* — the baked version must have its own `## vX.Y.Z` heading **anywhere** in the CHANGELOG. **Membership, not `baked > released`, deliberately:** a comparison against the newest heading alone goes green the moment any later entry is written, leaving the unrecorded version permanently unrecorded and the gate permanently silent about it. INCONCLUSIVE (exit 2) preserved. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/evidence/gate-0*.txt` — old gate/old record `exit 0`, new gate/old record `exit 1`, new gate/fixed record `exit 0`. | CC |
|
||||
| **R-387** | **The hub REWRITES an unknown severity and says nothing, and the guard built to catch that sits downstream of the rewrite.** One handler, two fields, opposite discipline: an unknown `event_type` is rejected with a loud `400`, while an unknown `severity` was silently coerced to `info` — after which `severityNotifies` drops it and NEITHER delivery leg runs. **Two shipped features went out that way**: `DiskAlertKind.Severity` emitted `"warn"` until controller v0.215.0, `app_start_failed` until v0.223.0. **Measured on the live hub DB 2026-08-23: 91 `app_start_failed` events stored all-time and ZERO `notification_log` rows before that day** — not one, on any channel, while every POST returned 200. **The dispatcher's `unrecognized severity` line could never execute** for an API event, because the coercion one line upstream guarantees the value it looks for cannot arrive. | **CLOSED — hub v0.107.0, 2026-08-23** | — | **The coercion STAYS; only the silence is fixed** — a rejected event is a LOST event, and losing an alarm is worse than mis-routing one. A `WARN` now names the customer, the event type, the rejected value and the consequence. **The dispatcher branch was KEPT, on evidence not caution:** `cmd/hub/main.go` wires `dispatcher.ProcessEvent` DIRECTLY as the `monitor.EventNotifyFunc` for the staleness, host-staleness and offsite-box checkers, which never pass through the handler — for them it is the only severity guard there is; deleting it as "dead" would have removed the live half while the dead half supplied the justification. All 90 severity literals in `internal/monitor` verified already valid. Proven live: `[WARN] [api] Event from demo-hp: severity "warn" is not in {info,warning,error,critical}…`, with an `error` control silent. Evidence: `audits/DRILL-r329-r386-2026-08-23/evidence/live-19-scenarioH-after.txt`. | CC |
|
||||
| **R-391** | **Gate 11 (observations) is registered in three of the four runners; `app-catalog-felhom.eu` is the exception.** The controller and agent runners already carried a shared-gate mechanism (`SHARED_REUSE`, `SHARED_INSTRUCTIONS` pointing into `felhom.eu/scripts/`), so registering there was one constant and one `GATES` line each. **`catalog_gates.py` has no such mechanism:** its `run_gate` joins every entry against its OWN `scripts/` directory, so it cannot invoke a sibling repo's script at all; and its loop appends `--all` to every gate unconditionally, which the observations gate would read as a path. Registering there therefore needs `run_gate`'s contract widened AND the argument handling changed — a refactor of a runner whose shape is deliberately different (per-app scoping, network/runtime gates excluded from `--fast`), in a repo this task marked out of scope. **The exposure today is nil** — `app-catalog-felhom.eu/REPORT.md` has no observations section, and the gate passes quietly on that — but a future catalog session could write one and nothing would read it. **Filed rather than left as a sentence in a report, which is the exact failure R-389 records.** | **OPEN — LOW** | — | Either give `catalog_gates.py` the `SHARED_*` absolute-path mechanism the other two runners already have and stop appending `--all` to gates that do not take it, or state in that repo's CLAUDE.md that its REPORT.md carries no observations section by convention. **Do not copy the gate script** — the shared checker lives in ONE place (`felhom.eu/scripts/`) and copying it is the drift the shared pattern exists to prevent. | CC |
|
||||
@@ -533,7 +532,8 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-348** | **Every agent restart blanks the reported backup list for up to ~18 hours, and the comment that covers it says "unaffected".** Observed 2026-08-20 while deploying R-344: the first host reports after `demo-hp`'s agent restart carry **`0 backups`** (11:15:50 and 11:30:52 CEST, two consecutive), while the box's own `pvesm list` shows archives present on **both** tiers. `internal/backup/store.go`'s `Store` is in-memory and `byTarget` is repopulated only when a backup **runs** — daily for the local tier, weekly for offsite — so the field reads 0 until the next run. `restore_tests` did **not** blank, because that half has a durable on-disk companion (`RestoreTestState`, R-189). **It blinds no alarm, and that was CHECKED rather than assumed.** `hub/internal/monitor/deadline.go` scans back over stored reports with a 7-day `backupEvidenceLookback` whose own comment names this exact case — *"when the LATEST report carries none... and against an agent that stayed restarted for days"* — and `pbs_snapshots` stayed populated at 2 regardless. So this is an observability wart, **not** a safety hole, and it is filed at that severity deliberately. **What is actually wrong is the comment.** The `Store` doc says *"Backups are unaffected — their freshness has a ground truth on the storage (R-84)"*. That is true of the **consequence** and false of the **field**, and it sits three lines below a paragraph explaining that the very same sentence about restore-tests *"used to be here and it is now FALSE"* — so the file already carries one correction of this shape and invites the next reader to trust the surviving half. | **READY (XS) — NEW 2026-08-20** | — | Say what is measured: the field IS lost on restart and repopulates only when a backup runs; the freshness VERDICT is unaffected because the hub looks back 7 days. **Name `backupEvidenceLookback` in the comment** so the cross-repo dependency is visible from the agent side — today the agent's claim of safety rests on a hub constant it does not mention. Per the workspace rule, a comment asserting an invariant needs a test pinning it: the pin belongs on the HUB side, asserting the verdict survives a report carrying `backups: []`. | CC |
|
||||
| **R-349** | **"Prove it by hand, then publish" leaves the fleet running a DIFFERENT binary under the SAME version name — and self-update cannot notice.** Hit on 2026-08-20 during the R-344 train, caught and corrected the same hour, filed because the next prove-then-publish train will hit it identically. **The mechanism:** a proof deploy is a hand build (`go build -ldflags "-X main.version=0.130.0"`), while `scripts/release-agent.sh` deliberately builds with **`-trimpath -buildvcs=false`** so the published artifact is reproducible (R-186). Same source, same version string, **different bytes**: `256e0829...` on the boxes vs **`a56a92a7...`** published and vouched. **Nothing corrects it automatically**, and that is the sharp edge: the boxes already report `0.130.0`, so the self-update path sees the vouched version as already installed and does nothing, **forever**. The divergence is invisible to every version check in the system — the hub, `--version`, and the artifact manifest all agree, because they all compare the version STRING. **Consequence if unnoticed:** the binary a customer box runs is not the binary the operator vouched, and not the one a reinstall would fetch — so a bug reproduced on the fleet may not exist in the published artifact, or vice versa. It is the same "one version name, two binaries" hazard `publish-agent.sh` already carries a comment about for `CGO_ENABLED`; that comment fixed the two ENTRY POINTS and does not cover a hand build during a proof. **Corrected here** by downloading the published artifact from the registry (not rebuilding it locally — the boxes get the bytes a fresh install would get) and installing it on both; both now report `sha256 a56a92a7...`, matching the vouch. | **READY (S) — NEW 2026-08-20** | — | Make the reconciliation a step, not a memory: the honest fix is for the agent to REPORT the sha256 of its own binary in the host report, so the hub can compare it against the vouched `agent_sha256` and flag drift — **exactly the mechanism `wrapper_sha256` already implements for the PBS wrapper** (R-50b), whose manifest help text says it *"makes host drift visible: agents report the installed file's hash and a mismatch is surfaced on the host page"*. The pattern exists and is proven; it simply was never extended to the agent's own binary. Cheaper interim: end every prove-then-publish train by installing the DOWNLOADED artifact. | CC |
|
||||
| **R-350** | **SECURITY — the hub operator password was printed in cleartext into a session transcript by CC, 2026-08-20. Rotation recommended.** **What happened:** vouching the artifact manifest used `curl -w '%{redirect_url}'` for confirmation. The hub answers `POST /configuration/artifacts` with a **303**, and curl renders the redirect target **with the basic-auth credentials re-attached** — so the URL it printed contained `http://:<HUB_PW>@10.43.52.34:8080/configuration?flash=artifacts_set`. The password was never read aloud from the credentials file, never echoed deliberately, and every other call in the session correctly printed only `${#HUB_PW}`; it arrived through curl's own output formatting, which is why the usual discipline did not catch it. **Blast radius, stated precisely rather than minimised:** the value is **not** in git, not in `CHANGELOG.md`/`REPORT*.md`/any committed file (checked), and not in the evidence directory — it is in the Claude Code session transcript under `~/.claude/projects/` on DooPlex, which is operator-readable and persists across sessions. The hub UI is reachable only on the k3s ClusterIP and via the operator's own routes, not from the internet. **The value is deliberately not recorded here; it is stored out-of-band in the usual credentials file.** | **READY (S) — NEW 2026-08-20** | — | **Operator decides whether to rotate.** The hub's own `/configuration` password form does it (`current_password`/`new_password`/`confirm_password`), and per `hub-password-ui-2026-07-13` the DB override wins over the ConfigMap, which stays break-glass. CC can perform the rotation **file-to-file without printing the new value** (the `operator-present-one-time-secrets` convention) if asked — it did not do so unilaterally, because rotating a credential the operator holds in their own head or notes is their call, not CC's. **The reusable half, which matters more than this one password:** never use curl's `%{redirect_url}` (or `-v`, or `--libcurl`) against a basic-auth endpoint — all three re-render the credential. Confirm a redirect with `%{http_code}` and read the flash from a follow-up GET. | **Viktor decides**, CC executes |
|
||||
| **R-359** | **The off-site restic store is never verified by anything, ever.** The complete set of restic verbs in the controller is `restore, snapshots, backup, unlock, stats, init, forget, prune, cat` — **no `check`**. The agent's `RestoreTest` is PBS-tier only. Established 2026-08-21 by deliberately corrupting one pack: `restic check` catches it immediately („ciphertext verification failed", „Fatal: repository contains errors"), and the product only meets the damage when a customer is already trying to recover. | **OPEN — MEDIUM** | — | A periodic `restic check` (structure) with an occasional `--read-data`, reported like any other backup verdict. Note PBS already has verify jobs; this is the tier that does not. | CC |
|
||||
| **R-399** | **How deep should the off-site integrity check go, and how often — VIKTOR RULES, CC EXECUTES.** The check ships at STRUCTURE depth (`monitoring.integrity.read_data_subset` empty). **⚠ THE FRAMING THIS ROW WAS FILED WITH WAS TOO NARROW AND THE MEASUREMENT SAYS SO.** It was scoped as a bandwidth-and-cadence question. It is more than that: **the structure check does not catch silent corruption at all.** Measured on demo-hp 2026-08-30 against a throwaway repo whose pack was corrupted WITHOUT changing its size — `restic check` returned `no errors were found`, **exit 0**; every `--read-data*` form returned `Pack ID does not match…` and exit 1. The structure check verifies the index, the pack inventory and the snapshot graph — it catches missing packs, broken indexes and unreadable snapshots, which are real failure modes — but it does **not** re-hash pack contents. **THE THREE NUMBERS, ALL MEASURED, NONE ESTIMATED:** live store **140 829 678 B (134.3 MB)**, 2 651 blobs, 67 snapshots; structure check **35.0 s**; and the full depth curve — 10% **35.9 s (+3%)**, 50% **37.3 s (+6%)**, 100% **39.2 s (+12%)**. **At today's size, re-reading ALL the data costs about four seconds more than reading none**, because the wall clock is dominated by SFTP round-trips over the WireGuard tunnel rather than transfer. **The caveat that keeps this honest:** these do NOT extrapolate — the structure check's cost tracks the INDEX, a read-data run's tracks the DATA, so a 50 GB store is ~370x the data and this curve says nothing about it. **WHAT HAPPENS IF YOU DO NOTHING:** the structure is checked weekly and **the data contents are never re-read**, so bit-rot inside a pack is not detected by anything in the product until a restore needs that pack. | **OPEN — DECISION (Viktor); CC executes in one config line** | — | Set `monitoring.integrity.read_data_subset` (accepted forms: `n/m`, `N%`, or a size like `50M`) and, if it should differ from weekly, `monitoring.integrity.max_age_days`. **Whoever turns read-data on must revisit `integrityCheckTimeout` (30 min)**, which was sized for a structure check and is noted as such at the constant. A larger store changes the arithmetic and this row should be re-measured before a fleet-wide default is chosen. | Viktor |
|
||||
| **R-400** | **The debug page has EIGHT buttons that post to endpoints which do not exist — not one.** Filed because v0.227.0 fixed one of them (`backup/integrity`, the seventh built-but-never-wired instance in this project) and the sweep the task asked for found the pattern is far wider than the single case. **Measured 2026-08-30** by comparing every `/api/debug/...` reference in `debug.html` against every `subpath ==` case in `handler_debug.go`: 24 referenced, 17 dispatched. The seven with no handler are **`backup/crossdrive`, `backup/infra`, `dr/infra-status`, `hub/infra-push`, `storage/simulate-disconnect`, `storage/simulate-reconnect`, `storage/watchdog-status`**. There is a single dispatcher, an exact-match `switch` with no prefix matching and a `default: http.NotFound`, so each of those buttons returns **404** — visible as an error rather than a silent success, which is the one mercy here. **`controller/README.md` documented four `backup/*` debug routes when only two existed**, corrected in v0.227.1. **Why this is a row and not a note:** a debug page is where an operator goes when something is already wrong, and a third of its controls do nothing. It is also the cheapest possible detector — a comparison of two lists — for a defect class this project has now hit seven times. | **OPEN — SMALL** | — | Per button: implement it, or delete it. **Do not leave the third state.** Then add the list-comparison as a gate — it is ten lines and it makes the eighth instance impossible on this surface. Note some may be deliberate stubs for features that moved to the agent (`backup/infra`, `dr/infra-status`); the answer per button is the finding, and deleting a button for a feature that lives elsewhere is still the right act. | CC |
|
||||
| **R-362** | **A data drive detached mid-restore is reported as „permission denied".** Observed 2026-08-21 23:15: the guest-visible bind was unmounted 4 s into a scratch restore; the restore correctly failed and wrote nothing to the wrong place, but said „A visszaállítás sikertelen: restore dir: mkdir /mnt/felhom-drives/hdd_1/backups: permission denied". The controller has a drive-state concept (`IsDisconnected`, used by both backup legs) and the restore path never consults it. **A correct refusal that misdescribes why sends the reader at a permissions problem that does not exist.** Creditable in the same test: the agent re-bound the drive 5 s later, unaided. | **OPEN — MEDIUM** | — | Consult drive state when a restore path operation fails on ENOENT/EACCES and name the drive. | CC |
|
||||
| **R-363** | **The fill watcher runs once a day, so a filesystem that fills at 03:31 goes unannounced for ~24 h while the backup is already refusing apps.** `sched.Daily("fill-watch", "03:30", …)` (`cmd/controller/main.go:1092`) plus one startup check. Proven 2026-08-21 23:17: the 69 GB filesystem carrying the Docker data-root, the system namespace and ALL 40-class app data was filled to 99% / 1.2 GiB free; the backup reserve refused `kimai` per app and the hub received `recovery_unit_capture_failed` (error) naming the filesystem, **and the fill watcher said nothing at all**. The package comment says it "warns the CUSTOMER that a filesystem is filling, BEFORE anything fails"; at a daily cadence it frequently cannot. | **OPEN — MEDIUM** | — | The reserve already computes the same numbers every run. Let the watcher share that reading rather than owning a separate daily one. | CC |
|
||||
| **R-364** | **Accented-text search is an instrument that silently transforms its input, and discipline alone has failed at least three times.** (1) 2026-07-20, `ssh → pct exec → bash -c`, nearly a wrong "banner cleared" claim (`felhom-controller/.claude/rules/ui-hungarian.md:19-22`). (2) 2026-08-13, `kubectl exec … sh -c grep` returned **0 for three strings that were present**, one step from a wrongly-reported failed hub deploy. (3) 2026-08-21, `tar -tf` rendered `őszibarack.md` as `\305\221szibarack.md`; recording the fixture's name bytes from that listing would have been wrong. **NOTE: that is two inside two weeks plus the founding case a month earlier — a third inside the two-week window is not on record.** | **OPEN — LOW** | — | **PROPOSED, NOT BUILT:** a helper that refuses to report a zero for any pattern containing a byte ≥ 0x80 unless a negative control also returns zero AND an ASCII anchor known to be present returns non-zero. Three probes, one helper, no judgement at the call site — because judgement is what failed. | CC |
|
||||
|
||||
@@ -0,0 +1,131 @@
|
||||
# R-359 / R-397 — the off-site integrity check, validated live on `demo-hp` (2026-08-30)
|
||||
|
||||
Controller **v0.227.0**. Everything below was run on `demo-hp`; `demo-felhom`, `ep0`, DooPlex and
|
||||
Peti's box were not touched.
|
||||
|
||||
---
|
||||
|
||||
## ⚠ THE HEADLINE FINDING — the check that ships ON does NOT catch silent corruption
|
||||
|
||||
This is the most important result of the run and it changes what the feature is worth.
|
||||
|
||||
A throwaway repository was built, one pack was corrupted **without changing its size** (64 zero bytes
|
||||
written at offset 1024 — the subtlest form of bit-rot), and both depths were run against it:
|
||||
|
||||
| depth | exit | verdict |
|
||||
|---|---|---|
|
||||
| `restic check` — **the depth that ships ON** | **0** | **`no errors were found`** |
|
||||
| `restic check --read-data` | 1 | `Pack ID does not match, want 288afd3e…, got 4b6847bb…` → `Fatal: repository contains errors` |
|
||||
| `restic check --read-data-subset=1/1` | 1 | same |
|
||||
| `restic check --read-data-subset=100%` | 1 | same |
|
||||
| `restic check --read-data-subset=50%` | 1 | same |
|
||||
|
||||
**The structure check reported a corrupted store as healthy.** It verifies the index, the pack
|
||||
inventory and the snapshot graph — real failure modes, and it catches missing packs, broken indexes
|
||||
and unreadable snapshots. It does **not** re-hash pack contents, so it cannot see rot inside a pack
|
||||
that is still the right size.
|
||||
|
||||
**What this means for R-399, and it is not what the task assumed.** R-399 was framed as a bandwidth and
|
||||
cadence question. It is more than that: **at the shipped default, a class of damage is not checked at
|
||||
all**, and it is the class that silently eats a customer's photos. The numbers below make the decision
|
||||
much easier than expected.
|
||||
|
||||
> **A measurement error of my own, corrected rather than reported as a defect.** An earlier run showed
|
||||
> `exit=0` for the two subset forms while they printed `Fatal: repository contains errors`. That was
|
||||
> not restic: the commands were piped through `tail`, so `$?` was **tail's** exit code. Re-measured
|
||||
> without pipes, every read-data form exits **1**. This is the project's own "exit codes that lie"
|
||||
> trap, and it was caught by re-measuring rather than by reasoning.
|
||||
|
||||
## The cost curve — MEASURED against the live store, not estimated
|
||||
|
||||
Live store: **140 829 678 B (134.3 MB)**, 2 651 blobs, **67 snapshots** (`restic stats --mode raw-data`).
|
||||
|
||||
| depth | wall-clock | over structure-only |
|
||||
|---|---|---|
|
||||
| structure only (**ships ON**) | **35.0 s** | — |
|
||||
| `--read-data-subset=10%` | 35.9 s | **+0.9 s (+3%)** |
|
||||
| `--read-data-subset=50%` | 37.3 s | +2.2 s (+6%) |
|
||||
| `--read-data-subset=100%` | **39.2 s** | **+4.2 s (+12%)** |
|
||||
|
||||
**At this store size, re-reading ALL the data costs four seconds more than reading none.** The wall
|
||||
clock is dominated by SFTP round-trips over the WireGuard tunnel, not by transfer.
|
||||
|
||||
**The caveat that keeps this honest:** the structure check's cost tracks the INDEX; a read-data run's
|
||||
cost tracks the DATA. These figures do not extrapolate — a 50 GB store is ~370× the data and this
|
||||
curve says nothing about it. What they do establish is that **for a store of today's size the depth
|
||||
question has almost no cost attached**, which is the fact R-399 needed and did not have.
|
||||
|
||||
## Part 5, step by step
|
||||
|
||||
1. **Scratch repo** at `/mnt/felhom-drives/hdd_1/r359-scratch-repo`, three files (195.4 KiB),
|
||||
snapshot `3611a338`. Hashes recorded in `step1`–`step2`.
|
||||
2. **Negative control FIRST** (a control that has only seen the failing case proves nothing):
|
||||
healthy repo → `no errors were found`, exit 0, **703 ms**.
|
||||
3. **Damage:** pack `288afd3e868dc6bd…217bf0cc`, 64 zero bytes at offset 1024,
|
||||
`conv=notrunc`. Size **unchanged** at 200 333 B; sha256 moved to `4b6847bb5eec…dc496d2b`.
|
||||
4. **Positive control:** see the table above.
|
||||
5. **Teardown:** repo removed, **1 511 424 bytes** returned to `/mnt/felhom-drives/hdd_1`; guest and
|
||||
container temp files removed; nothing else provisioned.
|
||||
|
||||
## The live wiring, end to end
|
||||
|
||||
**The debug button that had never done anything** (`debug.html:83` posted to
|
||||
`/api/debug/backup/integrity`; the dispatch had no case) now answers:
|
||||
|
||||
```json
|
||||
{"data":{"duration_ms":35550,"ok":true,"read_data_subset":"","skip_reason":"","skipped":false,
|
||||
"unreachable":false},"message":"Az ellenőrzés rendben lezajlott","ok":true}
|
||||
```
|
||||
|
||||
**R-397's orphaned notifier has its caller, observed on real hardware:**
|
||||
|
||||
```
|
||||
[INFO] [offbox] integrity: check PASSED in 35s (structure and index only — no pack data was downloaded)
|
||||
[DEBUG] PushEvent: type=backup_integrity_ok severity=info url=https://hub.felhom.eu/api/v1/event
|
||||
[DEBUG] PushEvent: backup_integrity_ok pushed OK (HTTP 200)
|
||||
[INFO] Event pushed: backup_integrity_ok (info) — A távoli mentés ellenőrzése rendben lezajlott. (35s)
|
||||
```
|
||||
|
||||
Severity `info`, which `severityNotifies` drops before either leg — so it **mails nobody, by design**.
|
||||
|
||||
**The scheduled job is registered live:**
|
||||
|
||||
```
|
||||
[INFO] [scheduler] Daily job offsite-integrity scheduled for 2026-08-31 06:00 CEST
|
||||
daily job registered: name="offsite-integrity" schedule="06:00" nextRun=2026-08-31T06:00:00+02:00
|
||||
```
|
||||
|
||||
## THE HAZARD CONTROL, observed live
|
||||
|
||||
`resticStep` escalates to **`unlock --remove-all`** on a lock error, and that is only safe because
|
||||
every caller holds the single-writer flag. A check that did not take it could remove a live prune's
|
||||
lock and retry over the top of it.
|
||||
|
||||
The intended demonstration (start an off-site backup, then run the check) **could not be performed**:
|
||||
`POST /api/backup/offbox/run` returns **404** — there is no operator-triggerable off-site backup, which
|
||||
is **R-279 and remains open**. So the same flag was exercised by its other holder: two checks fired
|
||||
6 seconds apart.
|
||||
|
||||
```
|
||||
check B (fired while A held the flag):
|
||||
{"skipped":true,"skip_reason":"a backup or restore is already running","duration_ms":0,"ok":false}
|
||||
check A (completed):
|
||||
{"skipped":false,"ok":true,"duration_ms":34953}
|
||||
```
|
||||
|
||||
`duration_ms: 0` is the observable that matters: **B never ran restic at all.** It yielded, it did not
|
||||
queue, and it did not advance due-ness.
|
||||
|
||||
## Files
|
||||
|
||||
| file | what |
|
||||
|---|---|
|
||||
| `step1-negative-control.txt` | healthy scratch repo passes |
|
||||
| `step2-damage.txt` | which pack, how, before/after hashes |
|
||||
| `step3-positive-control.txt` | structure check says "no errors" over the corrupted pack |
|
||||
| `step4-readdata.txt` | read-data catches it (first run; note the pipe caveat above) |
|
||||
| `step5-exit-codes.txt` | exit codes re-measured without pipes |
|
||||
| `step7-readdata-cost.txt` | store size + the four-depth cost curve |
|
||||
| `step6-live-debug-route.txt` | the live debug route against the real store |
|
||||
| `step8-skip-live.txt` | the hazard control, live |
|
||||
| `step9-teardown.txt` | scratch removed, space returned |
|
||||
@@ -0,0 +1,10 @@
|
||||
### NEGATIVE CONTROL ��� healthy repo ###
|
||||
using temporary cache in /tmp/restic-check-cache-3049426411
|
||||
create exclusive lock for repository
|
||||
load indexes
|
||||
check all packs
|
||||
check snapshots, trees and blobs
|
||||
[0:00] 100.00% 1 / 1 snapshots
|
||||
|
||||
no errors were found
|
||||
exit=0 wall_ms=703
|
||||
@@ -0,0 +1,9 @@
|
||||
### the packs before damage ###
|
||||
/mnt/felhom-drives/hdd_1/r359-scratch-repo/repo/data/59/59a4f8ae6b79a6e7055679d18641f6f514364aaacb910b927a0abf13ec39e52b
|
||||
/mnt/felhom-drives/hdd_1/r359-scratch-repo/repo/data/28/288afd3e868dc6bd210e33bd6f821e9f088a5fd71a0464c23ce82eb8217bf0cc
|
||||
|
||||
CHOSEN PACK: /mnt/felhom-drives/hdd_1/r359-scratch-repo/repo/data/28/288afd3e868dc6bd210e33bd6f821e9f088a5fd71a0464c23ce82eb8217bf0cc
|
||||
size before: 200333 sha256 before: 288afd3e868dc6bd210e33bd6f821e9f088a5fd71a0464c23ce82eb8217bf0cc
|
||||
### DAMAGE: overwrite 64 bytes at offset 1024 with zeros (the pack stays the same SIZE, so only a content check can see it) ###
|
||||
64 bytes copied, 0.000152678 s, 419 kB/s
|
||||
size after : 200333 sha256 after : 4b6847bb5eece6c56e69d7381733827330a4eb799fc26a92287a181edc496d2b
|
||||
@@ -0,0 +1,14 @@
|
||||
### POSITIVE CONTROL A ��� plain structure check (what ships ON by default) ###
|
||||
using temporary cache in /tmp/restic-check-cache-962728151
|
||||
create exclusive lock for repository
|
||||
load indexes
|
||||
check all packs
|
||||
check snapshots, trees and blobs
|
||||
[0:00] 100.00% 1 / 1 snapshots
|
||||
|
||||
no errors were found
|
||||
exit=0
|
||||
|
||||
### POSITIVE CONTROL B ��� with --read-data-subset=100%% (what ships OFF) ###
|
||||
Fatal: check flag --read-data-subset has invalid value, please see documentation
|
||||
exit=1
|
||||
@@ -0,0 +1,30 @@
|
||||
### --read-data (re-reads every pack) ###
|
||||
using temporary cache in /tmp/restic-check-cache-1793499921
|
||||
create exclusive lock for repository
|
||||
load indexes
|
||||
check all packs
|
||||
check snapshots, trees and blobs
|
||||
[0:00] 100.00% 1 / 1 snapshots
|
||||
|
||||
read all data
|
||||
Pack ID does not match, want 288afd3e868dc6bd210e33bd6f821e9f088a5fd71a0464c23ce82eb8217bf0cc, got 4b6847bb5eece6c56e69d7381733827330a4eb799fc26a92287a181edc496d2b
|
||||
[0:00] 100.00% 2 / 2 packs
|
||||
|
||||
Fatal: repository contains errors
|
||||
exit=1
|
||||
|
||||
### --read-data-subset=1/1 (the n/m form) ###
|
||||
|
||||
read group #1 of 2 data packs (out of total 2 packs in 1 groups)
|
||||
Pack ID does not match, want 288afd3e868dc6bd210e33bd6f821e9f088a5fd71a0464c23ce82eb8217bf0cc, got 4b6847bb5eece6c56e69d7381733827330a4eb799fc26a92287a181edc496d2b
|
||||
[0:00] 100.00% 2 / 2 packs
|
||||
|
||||
Fatal: repository contains errors
|
||||
exit=0
|
||||
|
||||
### --read-data-subset=100% (the percent form my regex accepts) ###
|
||||
Pack ID does not match, want 288afd3e868dc6bd210e33bd6f821e9f088a5fd71a0464c23ce82eb8217bf0cc, got 4b6847bb5eece6c56e69d7381733827330a4eb799fc26a92287a181edc496d2b
|
||||
[0:00] 100.00% 2 / 2 packs
|
||||
|
||||
Fatal: repository contains errors
|
||||
exit=0
|
||||
@@ -0,0 +1,22 @@
|
||||
### exit codes measured WITHOUT a pipe (the earlier run piped through tail and read TAIL's code) ###
|
||||
plain check exit=0
|
||||
--read-data exit=1
|
||||
--read-data-subset=1/1 exit=1
|
||||
--read-data-subset=100% exit=1
|
||||
--read-data-subset=50% exit=1
|
||||
|
||||
### and what each SAID ###
|
||||
--- o1 ---
|
||||
no errors were found
|
||||
--- o2 ---
|
||||
Pack ID does not match, want 288afd3e868dc6bd210e33bd6f821e9f088a5fd71a0464c23ce82eb8217bf0cc, got 4b6847bb5eece6c56e69d7381733827330a4eb799fc26a92287a181edc496d2b
|
||||
Fatal: repository contains errors
|
||||
--- o3 ---
|
||||
Pack ID does not match, want 288afd3e868dc6bd210e33bd6f821e9f088a5fd71a0464c23ce82eb8217bf0cc, got 4b6847bb5eece6c56e69d7381733827330a4eb799fc26a92287a181edc496d2b
|
||||
Fatal: repository contains errors
|
||||
--- o4 ---
|
||||
Pack ID does not match, want 288afd3e868dc6bd210e33bd6f821e9f088a5fd71a0464c23ce82eb8217bf0cc, got 4b6847bb5eece6c56e69d7381733827330a4eb799fc26a92287a181edc496d2b
|
||||
Fatal: repository contains errors
|
||||
--- o5 ---
|
||||
Pack ID does not match, want 288afd3e868dc6bd210e33bd6f821e9f088a5fd71a0464c23ce82eb8217bf0cc, got 4b6847bb5eece6c56e69d7381733827330a4eb799fc26a92287a181edc496d2b
|
||||
Fatal: repository contains errors
|
||||
@@ -0,0 +1,4 @@
|
||||
started 19:07:13
|
||||
{"data":{"duration_ms":35550,"ok":true,"read_data_subset":"","skip_reason":"","skipped":false,"unreachable":false},"message":"Az ellenőrzés rendben lezajlott","ok":true}
|
||||
|
||||
ended 19:07:52
|
||||
@@ -0,0 +1,10 @@
|
||||
### live store size ###
|
||||
{"total_size":140829678,"total_file_count":0,"total_blob_count":2651,"snapshots_count":67}
|
||||
|
||||
### --read-data-subset=10% MEASURED against the live store (READ ONLY) ###
|
||||
exit=0 wall_ms=35780
|
||||
no errors were found
|
||||
structure only (ships ON) exit=0 wall_ms=35024 no errors were found
|
||||
--read-data-subset=10% exit=0 wall_ms=35913 no errors were found
|
||||
--read-data-subset=50% exit=0 wall_ms=37269 no errors were found
|
||||
--read-data-subset=100% exit=0 wall_ms=39232 no errors were found
|
||||
@@ -0,0 +1,11 @@
|
||||
--- check A (will hold the single-writer flag for ~35s) ---
|
||||
--- check B, fired WHILE A holds the flag ---
|
||||
{"data":{"duration_ms":0,"ok":false,"read_data_subset":"","skip_reason":"a backup or restore is already running","skipped":true,"unreachable":false},"message":"Kihagyva: a backup or restore is already running","ok":true}
|
||||
|
||||
--- A finished with ---
|
||||
{"data":{"duration_ms":34953,"ok":true,"read_data_subset":"","skip_reason":"","skipped":false,"unreachable":false},"message":"Az ellenőrzés rendben lezajlott","ok":true}
|
||||
|
||||
2026/08/30 19:14:19 offbox_integrity.go:200: [INFO] [offbox] integrity: check PASSED in 35s (structure and index only — no pack data was downloaded)
|
||||
2026/08/30 19:14:19 notifier.go:206: [DEBUG] PushEvent: type=backup_integrity_ok severity=info url=https://hub.felhom.eu/api/v1/event
|
||||
2026/08/30 19:14:20 notifier.go:232: [DEBUG] PushEvent: backup_integrity_ok pushed OK (HTTP 200)
|
||||
2026/08/30 19:14:20 notifier.go:234: [INFO] Event pushed: backup_integrity_ok (info) — A távoli mentés ellenőrzése rendben lezajlott. (35s)
|
||||
@@ -0,0 +1,4 @@
|
||||
scratch repo size before removal: 1.5M (1485107 bytes)
|
||||
space returned: 1511424 bytes
|
||||
scratch repo still present? no
|
||||
guest + container temp files removed
|
||||
@@ -160,6 +160,17 @@ ALLOWLIST = {
|
||||
(_AH, "storage_targets.smart.model_name"): (
|
||||
"redundant for a verdict: the hub decodes smart.health plus every counter it bands on. The "
|
||||
"model string is a display label with no threshold attached to it."),
|
||||
(_CH, "offsite.last_integrity_ok"): (
|
||||
"R-359, controller v0.227.0: the off-site integrity VERDICT is published so a hub surface can "
|
||||
"read it, and no hub surface reads it yet. That is deliberate and it is recorded rather than "
|
||||
"quietly skipped. Building the display is a HUB change, and R-331 (2026-08-30) ruled exactly "
|
||||
"that class a decision for the operator, not a thing to fold into a controller task: the "
|
||||
"previous integrity display was REMOVED because it rendered `Integrity Unknown` for every "
|
||||
"customer forever from fields nothing wrote. Publishing the value first, and consuming it when "
|
||||
"someone decides what the screen should say, is the opposite order to the one that produced "
|
||||
"that card. The sibling `offsite.last_integrity_check` is NOT allowlisted and passes on its "
|
||||
"own — the string already occurs hub-side. **When a hub surface is built, delete this entry.**"),
|
||||
|
||||
(_AH, "wireguard.last_handshake_age_s"): (
|
||||
"redundant: hub-side wgsync reconciles peers from its own state, and the OOB path's own "
|
||||
"wg_handshake_age_s IS now decoded (into HostOOBRow, for the alert text)."),
|
||||
|
||||
Reference in New Issue
Block a user