docs: controller v0.228.0 — R-399/R-400 closed, R-401/R-402 filed
gates / gates (push) Failing after 19s
gates / gates (push) Failing after 19s
STATUS.md: header said 2026-08-23 over a 2026-08-30 body, and two "Waiting on you" items were both numbered 4 — both fixed. R-399 leaves that section (decided and shipped); the depth change is stated in plain words and the remaining items each say what happens if Viktor does nothing. 00-capability-map.md: the off-site verification row now carries its DEPTH, and its live citation is the 2026-08-31 run at 100%. The weekly firing at the new depth stays IMPLEMENTED, not PROVEN-LIVE. 07-backup-architecture.md §10.2: R-399 recorded closed, with the one sentence that stops it being turned back down — the structure check PASSED a size-preserving pack corruption. R-87 untouched and still OPEN. Register: R-399 and R-400 compressed into CLOSED-ITEMS.md with their reasoning kept and 300d7e8 named as the commit holding the originals. R-401 filed with a TRIGGER (the slow-check WARN firing) rather than a date. R-402 filed: the integrity verdict and its depth are on the wire and no hub surface reads either. OPEN 166 -> 165, CLOSED 148 -> 150. wire_contract_gate.py: offsite.last_integrity_depth allowlisted WITH ITS REASON beside its sibling last_integrity_ok, both to be deleted together when a hub surface is built (R-402).
This commit is contained in:
@@ -1,7 +1,8 @@
|
|||||||
# STATUS — what works, what's broken, what's next
|
# STATUS — what works, what's broken, what's next
|
||||||
|
|
||||||
**Updated 2026-08-23 — you now hear about EVERY broken app, not just the first one each hour. The
|
**Updated 2026-08-31 — the weekly off-site check now re-reads your actual data, not just the list of
|
||||||
hub deployed itself; nothing is waiting on you except the floor from the last release.**
|
it. A third of the debug page did nothing and no longer exists. One thing is waiting on you: the
|
||||||
|
golden and the floor for 0.228.0.**
|
||||||
|
|
||||||
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
|
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
|
||||||
> part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.**
|
> part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.**
|
||||||
@@ -12,29 +13,16 @@ hub deployed itself; nothing is waiting on you except the floor from the last re
|
|||||||
*This section is allowed to be longer than one screen, and each item says what happens if you do
|
*This section is allowed to be longer than one screen, and each item says what happens if you do
|
||||||
nothing.*
|
nothing.*
|
||||||
|
|
||||||
1. **How deep should the off-site check go?** (R-399). The box now checks its off-site store weekly.
|
1. **Bake and vouch a golden that carries 0.228.0, and raise the floor.** The fleet is on **0.227.1**.
|
||||||
The check it runs today reads the catalogue — it catches a missing or unreadable backup, and it does
|
Controller **0.228.0** is built, deployed to `demo-hp` and proven there, but a golden and the floor
|
||||||
**not** re-read the stored bytes, so it cannot see a file that has quietly rotted.
|
are yours to move. **If you do nothing:** a machine installed tomorrow gets 0.227.1, whose weekly
|
||||||
**What it costs to go deeper, measured on your own machine today, not guessed:**
|
check reads only the catalogue.
|
||||||
the shallow check takes **35.0 s**; re-reading **all** the data takes **39.2 s**. Four seconds more.
|
|
||||||
That is because the time goes on talking to the off-site box, not on moving data — and it will stop
|
|
||||||
being true as the store grows, so this is worth re-measuring, not deciding once forever.
|
|
||||||
**If you do nothing:** the catalogue is checked weekly and the stored bytes are never re-read.
|
|
||||||
I can turn it on with one setting whenever you say.
|
|
||||||
|
|
||||||
2. **Nothing about delivery — the golden train is current.** Golden **0.227.1** was baked, published,
|
2. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
||||||
round-trip verified, vouched, and the fleet floor raised to 0.227.1 on 2026-08-30. Both demo
|
|
||||||
machines run it; **`demo-felhom` got there by itself** and started the new off-site check on its own
|
|
||||||
schedule without anyone touching it. A machine installed today receives 0.227.1 and everything
|
|
||||||
shipped today. Evidence: `documentation/tests/golden-0.227.1-2026-08-30/`.
|
|
||||||
|
|
||||||
3. **Nothing else.** Everything in the releases of 2026-08-30 is a fix to code that ships in the
|
|
||||||
controller image; no customer action, no data migration, no credential change.
|
|
||||||
|
|
||||||
4. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
|
||||||
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
|
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
|
||||||
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
|
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
|
||||||
4. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
|
||||||
|
3. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
||||||
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
|
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
|
||||||
as it misled one by an hour.
|
as it misled one by an hour.
|
||||||
|
|
||||||
@@ -131,12 +119,14 @@ off. **`peti-felhom` is a real machine we have not heard from since 15 July** an
|
|||||||
|
|
||||||
## Broken, or knowingly incomplete
|
## Broken, or knowingly incomplete
|
||||||
|
|
||||||
- **The off-site store is only checked SHALLOWLY, and the deep check is switched off** (R-399).
|
- **We do not know what the deep check costs on a BIG store** (R-401). Since 0.228.0 the weekly check
|
||||||
This is the honest version of what shipped today. The box now checks its own off-site store about
|
re-reads **all** your stored data, not just the list of it. We had to: a copy was damaged in a way
|
||||||
once a week (R-359, controller 0.227.0) — and the check it runs reads the *catalogue* of the backups,
|
that left its size unchanged, and the old shallow check said „no errors were found". Only the deep
|
||||||
not the backups themselves. **We proved the difference:** a copy was damaged in a way that left its
|
check caught it. **The cost we measured was four seconds** — 35.0 s before, 39.2 s after — but that
|
||||||
size unchanged, and the shallow check said „no errors were found". Only the deep check caught it.
|
was on a 134 MB store, and it will not stay four seconds. So the box now tells us: any check that
|
||||||
The deep check is built and **off**, waiting for your decision — see item 2 under „Waiting on you".
|
takes longer than five minutes writes a warning naming this item. **If you do nothing:** every
|
||||||
|
machine re-reads its whole store every week, however large it grows, and the first person to notice
|
||||||
|
would be a customer whose upload is busy. The warning is there so that does not happen.
|
||||||
- **We ask the off-site box a question about once a second** (R-336) — ~85,000 a day for a box we
|
- **We ask the off-site box a question about once a second** (R-336) — ~85,000 a day for a box we
|
||||||
write to weekly. The leak that made this dangerous is fixed (R-344); the volume is not. The ceiling
|
write to weekly. The leak that made this dangerous is fixed (R-344); the volume is not. The ceiling
|
||||||
is **under a year** away on the corrected measurement, not two.
|
is **under a year** away on the corrected measurement, not two.
|
||||||
|
|||||||
@@ -152,7 +152,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
|||||||
| **Network storage (NAS) is browse + bulk-media only — it may NOT host an app's data namespace** | controller v0.187.0 | **PROVEN-LIVE** (2026-07-30) | `audits/R108-network-app-namespace-2026-07-30.md`. Same-box before/after on demo-felhom through the exact endpoint the UI invokes (`POST /api/storage/migrate-app`, authenticated + CSRF): **pre-fix v0.186.0** the target was never examined — both a NAS-shaped and an unregistered path passed straight into `MigrateApp` and failed only on the app name (409); **post-fix v0.187.0** both are refused 400 with a Hungarian reason, while a real local drive still reaches `MigrateApp` (409) proving the guard is not over-broad. On demo-hp (the box with a REGISTERED share) the network-specific refusal fires. Non-effect verified in the registry: no `migrated_to`, nothing decommissioned, app `HDD_PATH` unchanged, no `backups/` on the share | **Closes R-108 and UNBLOCKS D5** (`07-backup-architecture.md` §7.3, §10.1). The share-root `:rslave` FileBrowser bind is deliberately UNCHANGED — load-bearing for automount wake, and unscopable — verified byte-identical by diffing demo-hp's generated compose before/after. Fails closed: `/mnt/felhom-drives` holds both kinds, so an unregistered path under it is un-classifiable and refused. **Not exercised:** the deploy POST and decommission-migrate refusals are unit-tested (non-effect, nil `stackMgr`) but were NOT live-fired — only migrate-app was. `.fab`-export-onto-NAS remains open (→ R-126) |
|
| **Network storage (NAS) is browse + bulk-media only — it may NOT host an app's data namespace** | controller v0.187.0 | **PROVEN-LIVE** (2026-07-30) | `audits/R108-network-app-namespace-2026-07-30.md`. Same-box before/after on demo-felhom through the exact endpoint the UI invokes (`POST /api/storage/migrate-app`, authenticated + CSRF): **pre-fix v0.186.0** the target was never examined — both a NAS-shaped and an unregistered path passed straight into `MigrateApp` and failed only on the app name (409); **post-fix v0.187.0** both are refused 400 with a Hungarian reason, while a real local drive still reaches `MigrateApp` (409) proving the guard is not over-broad. On demo-hp (the box with a REGISTERED share) the network-specific refusal fires. Non-effect verified in the registry: no `migrated_to`, nothing decommissioned, app `HDD_PATH` unchanged, no `backups/` on the share | **Closes R-108 and UNBLOCKS D5** (`07-backup-architecture.md` §7.3, §10.1). The share-root `:rslave` FileBrowser bind is deliberately UNCHANGED — load-bearing for automount wake, and unscopable — verified byte-identical by diffing demo-hp's generated compose before/after. Fails closed: `/mnt/felhom-drives` holds both kinds, so an unregistered path under it is un-classifiable and refused. **Not exercised:** the deploy POST and decommission-migrate refusals are unit-tested (non-effect, nil `stackMgr`) but were NOT live-fired — only migrate-app was. `.fab`-export-onto-NAS remains open (→ R-126) |
|
||||||
| **A Tier-1/Tier-2 app restore works from the DRIVE ALONE — the recovery unit carries the app's secrets (D5)** | controller v0.188.0 | **PROVEN-LIVE** (2026-07-30) | `audits/D5-drive-alone-restore-2026-07-30.md`, `07-backup-architecture.md` §7.4. On a scratch drill guest on felhom-pve, through the real endpoints (`POST /api/stacks/{app}/deploy` → `POST /api/backup/run` → `POST /backup/restore`): AdventureLog (`SECRET_KEY` **data_key** + `DB_PASSWORD`) restored with the guest's `app.yaml` **moved aside** → `secrets recovered=2/2`, **27.6 s**, `Restore-from-unit completed`. **The observable is the DATA, not the exit code:** the app itself then read the seeded customer row **over TCP with its own credential** (`connected_as=adventurelog over_TCP=True`), 51 Django tables intact, and the discriminator held — the pre-backup row returned while a row added AFTER the backup was gone, so the volume tar was genuinely restored. The unit held **no `.sql` dump**, so the DB came back from the volume tar, which is exactly the case a regenerated password would have broken. **The withheld half is proven too:** Grafana's `type: password` admin login was live in its container and `ENC:` in the guest, yet appeared in **0 files** anywhere under the backup namespace, and the unit's app.yaml header names it as withheld | **Ruling (operator, 2026-07-30): `type: secret` travels, `type: password` NEVER does, minus the `nonPortableSecrets` code register (`vaultwarden/ADMIN_TOKEN`).** Plaintext on the drive, like the data — defensible **only because** the internet-reachable class is withheld; the two are coupled and must not be relaxed independently. **The brief's own proposal (data_key-only) was tested and rejected:** the flag is unreliable (4 encryption keys the catalog itself labels as such are unflagged → **R-127**) and a DB password is not resettable in practice (`POSTGRES_PASSWORD` is ignored once PGDATA is non-empty). Precedence: **the unit wins** over the guest, because the unit's secrets match the data being restored. The fail-closed data-key gate is UNCHANGED. **Not exercised live:** the withheld-class O4 regeneration on restore, and Tier-2's own cross-drive copy of a secret-bearing unit (both unit-tested only). Venue was a **fixture-class scratch guest**, correct per `runbooks/target-selection.md` since D5's claim is about restore CODE, not the install path |
|
| **A Tier-1/Tier-2 app restore works from the DRIVE ALONE — the recovery unit carries the app's secrets (D5)** | controller v0.188.0 | **PROVEN-LIVE** (2026-07-30) | `audits/D5-drive-alone-restore-2026-07-30.md`, `07-backup-architecture.md` §7.4. On a scratch drill guest on felhom-pve, through the real endpoints (`POST /api/stacks/{app}/deploy` → `POST /api/backup/run` → `POST /backup/restore`): AdventureLog (`SECRET_KEY` **data_key** + `DB_PASSWORD`) restored with the guest's `app.yaml` **moved aside** → `secrets recovered=2/2`, **27.6 s**, `Restore-from-unit completed`. **The observable is the DATA, not the exit code:** the app itself then read the seeded customer row **over TCP with its own credential** (`connected_as=adventurelog over_TCP=True`), 51 Django tables intact, and the discriminator held — the pre-backup row returned while a row added AFTER the backup was gone, so the volume tar was genuinely restored. The unit held **no `.sql` dump**, so the DB came back from the volume tar, which is exactly the case a regenerated password would have broken. **The withheld half is proven too:** Grafana's `type: password` admin login was live in its container and `ENC:` in the guest, yet appeared in **0 files** anywhere under the backup namespace, and the unit's app.yaml header names it as withheld | **Ruling (operator, 2026-07-30): `type: secret` travels, `type: password` NEVER does, minus the `nonPortableSecrets` code register (`vaultwarden/ADMIN_TOKEN`).** Plaintext on the drive, like the data — defensible **only because** the internet-reachable class is withheld; the two are coupled and must not be relaxed independently. **The brief's own proposal (data_key-only) was tested and rejected:** the flag is unreliable (4 encryption keys the catalog itself labels as such are unflagged → **R-127**) and a DB password is not resettable in practice (`POSTGRES_PASSWORD` is ignored once PGDATA is non-empty). Precedence: **the unit wins** over the guest, because the unit's secrets match the data being restored. The fail-closed data-key gate is UNCHANGED. **Not exercised live:** the withheld-class O4 regeneration on restore, and Tier-2's own cross-drive copy of a secret-bearing unit (both unit-tested only). Venue was a **fixture-class scratch guest**, correct per `runbooks/target-selection.md` since D5's claim is about restore CODE, not the install path |
|
||||||
| **A restore SAYS what it returned, and refuses what it cannot do** — the four restore-surface truth defects from the 2026-08-21 drill | controller **v0.226.0** (R-353, R-357, R-358, R-360, R-396) | **PROVEN-LIVE (2026-08-30) for three of the four; R-357 is IMPLEMENTED only** | `audits/evidence-r353-r360-live-2026-08-30/live-validation.txt`, controller `CHANGELOG.md` v0.226.0 + `REPORT.md`. Driven on `demo-hp` through the endpoints the UI invokes (no browser on DooPlex; the residual is client-side rendering). **R-353:** the sentence read off the customer's own wizard page — `A(z) opengist: 1 adatkötet visszaállítva — az alkalmazás újraindult.` with real counts (1 volume of 1 listed, 0 databases of 0 listed, and correctly no database clause). **R-360:** in the exact flag state that produced the bug (display flag true, concurrency flag false) the delete was refused and **a planted canary file survived**. **R-358/R-396:** a `mode=unit` restore wrote `{"schema":1,…,"full":false}` at mode 0600 with no `.tmp` left, and the gate logged `scratch holds a UNIT-ONLY restore … place-to-live stays closed` | **WHAT IS AND IS NOT CLAIMED, split deliberately.** **R-357 (the destructive restore's free-space gate) is IMPLEMENTED, NOT PROVEN-LIVE** — filling a real filesystem is a drill step, not a build step, so it rests on seam tests (`SetOffboxFreeFn`, `SetOffboxSizer`, and the new `SetOffboxLatestSnapshotFn`) whose central assertion is that `StopStack` was never called. **R-353's Scenario B — the "backup held only settings" sentence — was NOT reproduced live either**, and the reason is stated rather than glossed: no app on `demo-hp` still has a data-less unit (the drill's opengist has been recaptured and now lists one volume dump), and falsifying a manifest to produce it is the hand-set-state shortcut this project forbids. That branch is unit-proven only. **This row is about the MESSAGE and the REFUSALS, not the recovery mechanism** — `07-backup-architecture.md` §8 row 3 keeps its PROVEN status because the restore always did return what the unit held; what it could not do was say so |
|
| **A restore SAYS what it returned, and refuses what it cannot do** — the four restore-surface truth defects from the 2026-08-21 drill | controller **v0.226.0** (R-353, R-357, R-358, R-360, R-396) | **PROVEN-LIVE (2026-08-30) for three of the four; R-357 is IMPLEMENTED only** | `audits/evidence-r353-r360-live-2026-08-30/live-validation.txt`, controller `CHANGELOG.md` v0.226.0 + `REPORT.md`. Driven on `demo-hp` through the endpoints the UI invokes (no browser on DooPlex; the residual is client-side rendering). **R-353:** the sentence read off the customer's own wizard page — `A(z) opengist: 1 adatkötet visszaállítva — az alkalmazás újraindult.` with real counts (1 volume of 1 listed, 0 databases of 0 listed, and correctly no database clause). **R-360:** in the exact flag state that produced the bug (display flag true, concurrency flag false) the delete was refused and **a planted canary file survived**. **R-358/R-396:** a `mode=unit` restore wrote `{"schema":1,…,"full":false}` at mode 0600 with no `.tmp` left, and the gate logged `scratch holds a UNIT-ONLY restore … place-to-live stays closed` | **WHAT IS AND IS NOT CLAIMED, split deliberately.** **R-357 (the destructive restore's free-space gate) is IMPLEMENTED, NOT PROVEN-LIVE** — filling a real filesystem is a drill step, not a build step, so it rests on seam tests (`SetOffboxFreeFn`, `SetOffboxSizer`, and the new `SetOffboxLatestSnapshotFn`) whose central assertion is that `StopStack` was never called. **R-353's Scenario B — the "backup held only settings" sentence — was NOT reproduced live either**, and the reason is stated rather than glossed: no app on `demo-hp` still has a data-less unit (the drill's opengist has been recaptured and now lists one volume dump), and falsifying a manifest to produce it is the hand-set-state shortcut this project forbids. That branch is unit-proven only. **This row is about the MESSAGE and the REFUSALS, not the recovery mechanism** — `07-backup-architecture.md` §8 row 3 keeps its PROVEN status because the restore always did return what the unit held; what it could not do was say so |
|
||||||
| **The off-site store is VERIFIED on a cadence — something checks that the customer's backups are still readable** | controller **v0.227.0/v0.227.1** (R-359, R-397) | **PROVEN-LIVE (2026-08-30) for the check, the notifier and the hazard control; the SCHEDULED FIRING is IMPLEMENTED only** | `documentation/tests/r359-integrity-2026-08-30/`. Driven on `demo-hp` through the endpoint the debug button invokes. A throwaway repo was built, checked healthy (**negative control first**), then one pack corrupted; the live store was checked read-only in **35.0 s**; and the notifier fired end to end — `Event pushed: backup_integrity_ok (info)`. The hazard control was observed live: a second check fired while the first held the single-writer flag returned `skipped:true, duration_ms:0` — **it never ran restic at all** | **⚠ WHAT AN `ok` DOES AND DOES NOT MEAN, and this is the row's most important sentence.** The depth that ships ON is a STRUCTURE check: index, pack inventory, snapshot graph. **It does not re-hash pack contents, and it does not catch silent corruption** — measured, a pack corrupted without a size change returned `no errors were found`, exit 0, while every `--read-data*` form caught it. That is **R-399**, open, Viktor's decision, with the cost curve measured (structure 35.0 s · 10% 35.9 s · 50% 37.3 s · 100% 39.2 s on 134.3 MB — and those do NOT extrapolate). **The weekly firing is NOT proven-live** — it is a week away and this session could not observe it; the job is confirmed REGISTERED on the box (`Daily job offsite-integrity scheduled for 2026-08-31 06:00 CEST`), which is not the same claim. **This is a readability check and NOT a restore-test** — R-87 remains open and the two are routinely conflated because their register rows are adjacent |
|
| **The off-site store is VERIFIED on a cadence — something checks that the customer's backups are still readable** | controller **v0.228.0** (R-359, R-397, R-399) | **PROVEN-LIVE (2026-08-30, re-proven at FULL DEPTH 2026-08-31) for the check, the notifier and the hazard control; the SCHEDULED FIRING is IMPLEMENTED only** | `documentation/tests/r359-integrity-2026-08-30/`. Driven on `demo-hp` through the endpoint the debug button invokes. A throwaway repo was built, checked healthy (**negative control first**), then one pack corrupted; the live store was checked read-only in **35.0 s**; and the notifier fired end to end — `Event pushed: backup_integrity_ok (info)`. The hazard control was observed live: a second check fired while the first held the single-writer flag returned `skipped:true, duration_ms:0` — **it never ran restic at all** | **⚠ WHAT AN `ok` MEANS — CHANGED 2026-08-31 (R-399, controller v0.228.0): the check now RE-READS THE DATA.** The default is `--read-data-subset=100%`, so an `ok` means every stored byte was downloaded and re-hashed, not merely that the catalogue hangs together. **The reason is measured, and it is why the default must not be turned back down to save four seconds:** a pack corrupted WITHOUT a size change made a structure check return `no errors were found`, exit 0, while every `--read-data*` form caught it. Cost curve on 134.3 MB: structure 35.0 s · 10% 35.9 s · 50% 37.3 s · 100% 39.2 s — **and those do NOT extrapolate**, which is why v0.228.0 ships a slow-check WARN (R-401) rather than a rotation schedule. `off` returns a box to structure depth. **PROVEN-LIVE at the new depth 2026-08-31 on `demo-hp`**, endpoint-level, with the restic argv observed from the guest: default → `… check --read-data-subset=100%`, 38.7 s; `off` → `… check`, 34.7 s. **The weekly firing at the new depth is IMPLEMENTED only** — the job is confirmed REGISTERED on the box (`Daily job offsite-integrity scheduled for 2026-09-01 06:00 CEST`), which is not the same claim. **This is a readability check and NOT a restore-test** — R-87 remains open and the two are routinely conflated because their register rows are adjacent |
|
||||||
| USB drive enrollment + unplug detection + recommission | controller, agent | **PROVEN-LIVE** | `storage-lifecycle-acceptance-2026-06-15` E4 (yanked-while-running → agent auto-rebind) + E10 (re-enroll, data intact); `CAMPAIGN-4`/`6A` (3 USB re-establish across device-letter reshuffle) | (Cited `RUNBOOK-usb` could NOT complete a wizard enrollment; `CAMPAIGN-2` legs were auth-hollow.) Fresh-USB **wizard enrollment** specifically still unproven |
|
| USB drive enrollment + unplug detection + recommission | controller, agent | **PROVEN-LIVE** | `storage-lifecycle-acceptance-2026-06-15` E4 (yanked-while-running → agent auto-rebind) + E10 (re-enroll, data intact); `CAMPAIGN-4`/`6A` (3 USB re-establish across device-letter reshuffle) | (Cited `RUNBOOK-usb` could NOT complete a wizard enrollment; `CAMPAIGN-2` legs were auth-hollow.) Fresh-USB **wizard enrollment** specifically still unproven |
|
||||||
| Decommission (migrate-first and anyway-paths), eject | agent, controller | **PROVEN-LIVE** | `storage-lifecycle-acceptance-2026-06-15` E9 (decommission-anyway → bind detached, parent mp untouched, reboot-safe) + E12 (eject drive holding all apps) | (Cited `CAMPAIGN-2` T-STG-DECOM-* were auth-hollow; `SPIKE-decommission` was report-only, button still vestigial.) |
|
| Decommission (migrate-first and anyway-paths), eject | agent, controller | **PROVEN-LIVE** | `storage-lifecycle-acceptance-2026-06-15` E9 (decommission-anyway → bind detached, parent mp untouched, reboot-safe) + E12 (eject drive holding all apps) | (Cited `CAMPAIGN-2` T-STG-DECOM-* were auth-hollow; `SPIKE-decommission` was report-only, button still vestigial.) |
|
||||||
| Boot ordering: automount + networking survive reboot; appliance self-heal watchdog | agent v0.85 | **PROVEN-LIVE** | `CAMPAIGN-4-2026-07-13` (F12 fix HOLDS: demo-host reboot + 5-boot storm, 0 ordering cycles, caps 63/63, WG re-handshake) + `CAMPAIGN-6A-2026-07-14` 1D (re-arm reboot-survival across 9 guest + 1 host reboots) | (`CAMPAIGN-3` F10/F11/F12 were the CRITICAL/HIGH *failures*; fixes shipped in agent v0.85 and were re-validated live in 4/6A — cite the validation, not the finding.) Residual: `skip-active` on `pct reboot` carried by the heal path; a NAS outage spanning a guest reboot can strand the share until agent restart (6A) |
|
| Boot ordering: automount + networking survive reboot; appliance self-heal watchdog | agent v0.85 | **PROVEN-LIVE** | `CAMPAIGN-4-2026-07-13` (F12 fix HOLDS: demo-host reboot + 5-boot storm, 0 ordering cycles, caps 63/63, WG re-handshake) + `CAMPAIGN-6A-2026-07-14` 1D (re-arm reboot-survival across 9 guest + 1 host reboots) | (`CAMPAIGN-3` F10/F11/F12 were the CRITICAL/HIGH *failures*; fixes shipped in agent v0.85 and were re-validated live in 4/6A — cite the validation, not the finding.) Residual: `skip-active` on `pct reboot` carried by the heal path; a NAS outage spanning a guest reboot can strand the share until agent restart (6A) |
|
||||||
|
|||||||
@@ -1000,6 +1000,7 @@ does **not** hold as written. → **R-108**
|
|||||||
| ~~R-396~~ | ~~A unit-only verification restore unlocked the destructive full restore~~ | **CLOSED 2026-08-30, controller v0.226.0** (by R-358's marker). Found while answering R-358's open question. Both restore modes write the SAME scratch directory, and one boolean (`ScratchReady`) drove three different intents — "is there a scratch", "may we place", "may we destructively restore". The safest action on the page unlocked the most dangerous one |
|
| ~~R-396~~ | ~~A unit-only verification restore unlocked the destructive full restore~~ | **CLOSED 2026-08-30, controller v0.226.0** (by R-358's marker). Found while answering R-358's open question. Both restore modes write the SAME scratch directory, and one boolean (`ScratchReady`) drove three different intents — "is there a scratch", "may we place", "may we destructively restore". The safest action on the page unlocked the most dangerous one |
|
||||||
| ~~R-359~~ | ~~The off-site restic store is never verified by anything, ever~~ | **CLOSED 2026-08-30, controller v0.227.0/v0.227.1.** A daily `offsite-integrity` job on **due-ness, not a weekday**; it takes the single-writer flag and SKIPS rather than waits (`resticStep` escalates to `unlock --remove-all` and is only safe while that flag is held). Three outcomes — skipped / unreachable / failed — because 'I could not look' is not 'I looked and it is broken'. **⚠ The depth that ships ON does NOT catch silent corruption:** measured, a pack corrupted without a size change returned `no errors were found`, exit 0; only `--read-data*` caught it. Choosing the depth is **R-399** |
|
| ~~R-359~~ | ~~The off-site restic store is never verified by anything, ever~~ | **CLOSED 2026-08-30, controller v0.227.0/v0.227.1.** A daily `offsite-integrity` job on **due-ness, not a weekday**; it takes the single-writer flag and SKIPS rather than waits (`resticStep` escalates to `unlock --remove-all` and is only safe while that flag is held). Three outcomes — skipped / unreachable / failed — because 'I could not look' is not 'I looked and it is broken'. **⚠ The depth that ships ON does NOT catch silent corruption:** measured, a pack corrupted without a size change returned `no errors were found`, exit 0; only `--read-data*` caught it. Choosing the depth is **R-399** |
|
||||||
| ~~R-397~~ | ~~`NotifyIntegrityOK`/`NotifyIntegrityFailed` had no caller and the product advertised a weekly check that did not exist~~ | **CLOSED 2026-08-30, controller v0.227.0.** Sixth built-but-never-wired instance: hub allowlist, Hungarian text, settings checkbox and debug button all existed; only the caller did not. `ok` is severity `info` and mails nobody by design |
|
| ~~R-397~~ | ~~`NotifyIntegrityOK`/`NotifyIntegrityFailed` had no caller and the product advertised a weekly check that did not exist~~ | **CLOSED 2026-08-30, controller v0.227.0.** Sixth built-but-never-wired instance: hub allowlist, Hungarian text, settings checkbox and debug button all existed; only the caller did not. `ok` is severity `info` and mails nobody by design |
|
||||||
|
| ~~R-399~~ | ~~The check reads the catalogue and never the data~~ | **CLOSED 2026-08-31, controller v0.228.0.** `monitoring.integrity.read_data_subset` now defaults to **`100%`**, so the weekly check downloads and re-hashes every stored byte. **The fact that made it necessary, and the sentence that should stop anyone turning it back down to save four seconds: the structure check PASSED a size-preserving pack corruption.** Measured on `demo-hp` 2026-08-30 — plain `restic check` reported `no errors were found` and exited 0 over a pack damaged without a size change; every read-data form caught it. Cost on that 134 MB store: 35.0 s structure vs 39.2 s at 100%. `off` (any case) returns a box to structure depth; an empty value means *not configured*, therefore the default; a malformed value WARNs and falls back to the DEFAULT, never to structure. A completed check over 5 minutes logs an operator WARN naming R-401 — **one data point, on one 134 MB store, so no rotation schedule, size threshold or bandwidth budget was invented from it.** Proven live at both depths 2026-08-31 with the restic argv observed from the guest |
|
||||||
| **R-87 (open) — AND IT IS NOT R-359** | The restic tier is never restore-TESTED | matrix row 4's route has no unattended proof. **Stated explicitly because the two rows sit next to each other and are easy to conflate: a `check` proves the STORE IS READABLE; a restore-test proves DATA COMES BACK OUT.** v0.227.0 did the first and nothing else. This row is untouched by it |
|
| **R-87 (open) — AND IT IS NOT R-359** | The restic tier is never restore-TESTED | matrix row 4's route has no unattended proof. **Stated explicitly because the two rows sit next to each other and are easy to conflate: a `check` proves the STORE IS READABLE; a restore-test proves DATA COMES BACK OUT.** v0.227.0 did the first and nothing else. This row is untouched by it |
|
||||||
|
|
||||||
### 10.3 Divergences that are documented elsewhere and are not re-opened here
|
### 10.3 Divergences that are documented elsewhere and are not re-opened here
|
||||||
|
|||||||
@@ -233,4 +233,5 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
|
|||||||
hub, never in a doc.
|
hub, never in a doc.
|
||||||
- **A doc comment claiming a guard exists is why nobody looks for the missing guard** (R-360). Correct
|
- **A doc comment claiming a guard exists is why nobody looks for the missing guard** (R-360). Correct
|
||||||
such a sentence in place; do not delete it.
|
such a sentence in place; do not delete it.
|
||||||
|
| **R-399** | **How deep should the off-site integrity check go — Viktor ruled full depth.** Shipped in controller **v0.228.0**, 2026-08-31. Evidence: `felhom-controller/REPORT.md` (v0.228.0) — restic argv observed from the guest at both depths on `demo-hp`. **Reasoning kept:** *the structure check does not detect a size-preserving pack corruption — measured 2026-08-30, plain `restic check` reported `no errors were found` and exited 0 over a damaged pack that every read-data form caught. That is the reason for the default and it is what should stop anyone turning it back down to save four seconds.* *An empty value means "not configured", therefore the default; `off` is the off token, because a setting with no off switch is not a setting.* *A malformed value falls back to the DEFAULT, never to structure — falling back to structure would silently remove the protection on a typo, which is R-357's shape.* **Superseded by R-401** for anything about a large store. Original text: `git show 300d7e8:documentation/backlog/OPEN-ITEMS.md` |
|
||||||
|
| **R-400** | **A third of the debug page posted to endpoints that did not exist — and three of the seven fetched on page LOAD.** Shipped in controller **v0.228.0**, 2026-08-31. 24 referenced / 17 dispatched became 18 / 18. `backup/crossdrive` implemented (proven live: real Tier-2 copies for three apps); `backup/infra`, `hub/infra-push`, `dr/infra-status`, `storage/watchdog-status` and both `storage/simulate-*` deleted with their panels and JavaScript. **Reasoning kept:** *implement or delete FIRST, register the gate SECOND — a registered-but-failing gate refuses every push.* *Keep `handleDebugAPI`'s exact-match switch with its `NotFound` default; a prefix match would have made the defect invisible instead of merely silent.* *A panel left behind renders nothing forever, which is how this class hides.* *A debug control that simulates or mutates storage state is deleted unless a live need can be shown — that is where drives get unenrolled and data gets stranded.* Enforced by `controller/scripts/debug_route_gate.py`, both directions, red-proofed. Original text: `git show 300d7e8:documentation/backlog/OPEN-ITEMS.md` |
|
||||||
|
|||||||
@@ -532,8 +532,8 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
|||||||
| **R-348** | **Every agent restart blanks the reported backup list for up to ~18 hours, and the comment that covers it says "unaffected".** Observed 2026-08-20 while deploying R-344: the first host reports after `demo-hp`'s agent restart carry **`0 backups`** (11:15:50 and 11:30:52 CEST, two consecutive), while the box's own `pvesm list` shows archives present on **both** tiers. `internal/backup/store.go`'s `Store` is in-memory and `byTarget` is repopulated only when a backup **runs** — daily for the local tier, weekly for offsite — so the field reads 0 until the next run. `restore_tests` did **not** blank, because that half has a durable on-disk companion (`RestoreTestState`, R-189). **It blinds no alarm, and that was CHECKED rather than assumed.** `hub/internal/monitor/deadline.go` scans back over stored reports with a 7-day `backupEvidenceLookback` whose own comment names this exact case — *"when the LATEST report carries none... and against an agent that stayed restarted for days"* — and `pbs_snapshots` stayed populated at 2 regardless. So this is an observability wart, **not** a safety hole, and it is filed at that severity deliberately. **What is actually wrong is the comment.** The `Store` doc says *"Backups are unaffected — their freshness has a ground truth on the storage (R-84)"*. That is true of the **consequence** and false of the **field**, and it sits three lines below a paragraph explaining that the very same sentence about restore-tests *"used to be here and it is now FALSE"* — so the file already carries one correction of this shape and invites the next reader to trust the surviving half. | **READY (XS) — NEW 2026-08-20** | — | Say what is measured: the field IS lost on restart and repopulates only when a backup runs; the freshness VERDICT is unaffected because the hub looks back 7 days. **Name `backupEvidenceLookback` in the comment** so the cross-repo dependency is visible from the agent side — today the agent's claim of safety rests on a hub constant it does not mention. Per the workspace rule, a comment asserting an invariant needs a test pinning it: the pin belongs on the HUB side, asserting the verdict survives a report carrying `backups: []`. | CC |
|
| **R-348** | **Every agent restart blanks the reported backup list for up to ~18 hours, and the comment that covers it says "unaffected".** Observed 2026-08-20 while deploying R-344: the first host reports after `demo-hp`'s agent restart carry **`0 backups`** (11:15:50 and 11:30:52 CEST, two consecutive), while the box's own `pvesm list` shows archives present on **both** tiers. `internal/backup/store.go`'s `Store` is in-memory and `byTarget` is repopulated only when a backup **runs** — daily for the local tier, weekly for offsite — so the field reads 0 until the next run. `restore_tests` did **not** blank, because that half has a durable on-disk companion (`RestoreTestState`, R-189). **It blinds no alarm, and that was CHECKED rather than assumed.** `hub/internal/monitor/deadline.go` scans back over stored reports with a 7-day `backupEvidenceLookback` whose own comment names this exact case — *"when the LATEST report carries none... and against an agent that stayed restarted for days"* — and `pbs_snapshots` stayed populated at 2 regardless. So this is an observability wart, **not** a safety hole, and it is filed at that severity deliberately. **What is actually wrong is the comment.** The `Store` doc says *"Backups are unaffected — their freshness has a ground truth on the storage (R-84)"*. That is true of the **consequence** and false of the **field**, and it sits three lines below a paragraph explaining that the very same sentence about restore-tests *"used to be here and it is now FALSE"* — so the file already carries one correction of this shape and invites the next reader to trust the surviving half. | **READY (XS) — NEW 2026-08-20** | — | Say what is measured: the field IS lost on restart and repopulates only when a backup runs; the freshness VERDICT is unaffected because the hub looks back 7 days. **Name `backupEvidenceLookback` in the comment** so the cross-repo dependency is visible from the agent side — today the agent's claim of safety rests on a hub constant it does not mention. Per the workspace rule, a comment asserting an invariant needs a test pinning it: the pin belongs on the HUB side, asserting the verdict survives a report carrying `backups: []`. | CC |
|
||||||
| **R-349** | **"Prove it by hand, then publish" leaves the fleet running a DIFFERENT binary under the SAME version name — and self-update cannot notice.** Hit on 2026-08-20 during the R-344 train, caught and corrected the same hour, filed because the next prove-then-publish train will hit it identically. **The mechanism:** a proof deploy is a hand build (`go build -ldflags "-X main.version=0.130.0"`), while `scripts/release-agent.sh` deliberately builds with **`-trimpath -buildvcs=false`** so the published artifact is reproducible (R-186). Same source, same version string, **different bytes**: `256e0829...` on the boxes vs **`a56a92a7...`** published and vouched. **Nothing corrects it automatically**, and that is the sharp edge: the boxes already report `0.130.0`, so the self-update path sees the vouched version as already installed and does nothing, **forever**. The divergence is invisible to every version check in the system — the hub, `--version`, and the artifact manifest all agree, because they all compare the version STRING. **Consequence if unnoticed:** the binary a customer box runs is not the binary the operator vouched, and not the one a reinstall would fetch — so a bug reproduced on the fleet may not exist in the published artifact, or vice versa. It is the same "one version name, two binaries" hazard `publish-agent.sh` already carries a comment about for `CGO_ENABLED`; that comment fixed the two ENTRY POINTS and does not cover a hand build during a proof. **Corrected here** by downloading the published artifact from the registry (not rebuilding it locally — the boxes get the bytes a fresh install would get) and installing it on both; both now report `sha256 a56a92a7...`, matching the vouch. | **READY (S) — NEW 2026-08-20** | — | Make the reconciliation a step, not a memory: the honest fix is for the agent to REPORT the sha256 of its own binary in the host report, so the hub can compare it against the vouched `agent_sha256` and flag drift — **exactly the mechanism `wrapper_sha256` already implements for the PBS wrapper** (R-50b), whose manifest help text says it *"makes host drift visible: agents report the installed file's hash and a mismatch is surfaced on the host page"*. The pattern exists and is proven; it simply was never extended to the agent's own binary. Cheaper interim: end every prove-then-publish train by installing the DOWNLOADED artifact. | CC |
|
| **R-349** | **"Prove it by hand, then publish" leaves the fleet running a DIFFERENT binary under the SAME version name — and self-update cannot notice.** Hit on 2026-08-20 during the R-344 train, caught and corrected the same hour, filed because the next prove-then-publish train will hit it identically. **The mechanism:** a proof deploy is a hand build (`go build -ldflags "-X main.version=0.130.0"`), while `scripts/release-agent.sh` deliberately builds with **`-trimpath -buildvcs=false`** so the published artifact is reproducible (R-186). Same source, same version string, **different bytes**: `256e0829...` on the boxes vs **`a56a92a7...`** published and vouched. **Nothing corrects it automatically**, and that is the sharp edge: the boxes already report `0.130.0`, so the self-update path sees the vouched version as already installed and does nothing, **forever**. The divergence is invisible to every version check in the system — the hub, `--version`, and the artifact manifest all agree, because they all compare the version STRING. **Consequence if unnoticed:** the binary a customer box runs is not the binary the operator vouched, and not the one a reinstall would fetch — so a bug reproduced on the fleet may not exist in the published artifact, or vice versa. It is the same "one version name, two binaries" hazard `publish-agent.sh` already carries a comment about for `CGO_ENABLED`; that comment fixed the two ENTRY POINTS and does not cover a hand build during a proof. **Corrected here** by downloading the published artifact from the registry (not rebuilding it locally — the boxes get the bytes a fresh install would get) and installing it on both; both now report `sha256 a56a92a7...`, matching the vouch. | **READY (S) — NEW 2026-08-20** | — | Make the reconciliation a step, not a memory: the honest fix is for the agent to REPORT the sha256 of its own binary in the host report, so the hub can compare it against the vouched `agent_sha256` and flag drift — **exactly the mechanism `wrapper_sha256` already implements for the PBS wrapper** (R-50b), whose manifest help text says it *"makes host drift visible: agents report the installed file's hash and a mismatch is surfaced on the host page"*. The pattern exists and is proven; it simply was never extended to the agent's own binary. Cheaper interim: end every prove-then-publish train by installing the DOWNLOADED artifact. | CC |
|
||||||
| **R-350** | **SECURITY — the hub operator password was printed in cleartext into a session transcript by CC, 2026-08-20. Rotation recommended.** **What happened:** vouching the artifact manifest used `curl -w '%{redirect_url}'` for confirmation. The hub answers `POST /configuration/artifacts` with a **303**, and curl renders the redirect target **with the basic-auth credentials re-attached** — so the URL it printed contained `http://:<HUB_PW>@10.43.52.34:8080/configuration?flash=artifacts_set`. The password was never read aloud from the credentials file, never echoed deliberately, and every other call in the session correctly printed only `${#HUB_PW}`; it arrived through curl's own output formatting, which is why the usual discipline did not catch it. **Blast radius, stated precisely rather than minimised:** the value is **not** in git, not in `CHANGELOG.md`/`REPORT*.md`/any committed file (checked), and not in the evidence directory — it is in the Claude Code session transcript under `~/.claude/projects/` on DooPlex, which is operator-readable and persists across sessions. The hub UI is reachable only on the k3s ClusterIP and via the operator's own routes, not from the internet. **The value is deliberately not recorded here; it is stored out-of-band in the usual credentials file.** | **READY (S) — NEW 2026-08-20** | — | **Operator decides whether to rotate.** The hub's own `/configuration` password form does it (`current_password`/`new_password`/`confirm_password`), and per `hub-password-ui-2026-07-13` the DB override wins over the ConfigMap, which stays break-glass. CC can perform the rotation **file-to-file without printing the new value** (the `operator-present-one-time-secrets` convention) if asked — it did not do so unilaterally, because rotating a credential the operator holds in their own head or notes is their call, not CC's. **The reusable half, which matters more than this one password:** never use curl's `%{redirect_url}` (or `-v`, or `--libcurl`) against a basic-auth endpoint — all three re-render the credential. Confirm a redirect with `%{http_code}` and read the flash from a follow-up GET. | **Viktor decides**, CC executes |
|
| **R-350** | **SECURITY — the hub operator password was printed in cleartext into a session transcript by CC, 2026-08-20. Rotation recommended.** **What happened:** vouching the artifact manifest used `curl -w '%{redirect_url}'` for confirmation. The hub answers `POST /configuration/artifacts` with a **303**, and curl renders the redirect target **with the basic-auth credentials re-attached** — so the URL it printed contained `http://:<HUB_PW>@10.43.52.34:8080/configuration?flash=artifacts_set`. The password was never read aloud from the credentials file, never echoed deliberately, and every other call in the session correctly printed only `${#HUB_PW}`; it arrived through curl's own output formatting, which is why the usual discipline did not catch it. **Blast radius, stated precisely rather than minimised:** the value is **not** in git, not in `CHANGELOG.md`/`REPORT*.md`/any committed file (checked), and not in the evidence directory — it is in the Claude Code session transcript under `~/.claude/projects/` on DooPlex, which is operator-readable and persists across sessions. The hub UI is reachable only on the k3s ClusterIP and via the operator's own routes, not from the internet. **The value is deliberately not recorded here; it is stored out-of-band in the usual credentials file.** | **READY (S) — NEW 2026-08-20** | — | **Operator decides whether to rotate.** The hub's own `/configuration` password form does it (`current_password`/`new_password`/`confirm_password`), and per `hub-password-ui-2026-07-13` the DB override wins over the ConfigMap, which stays break-glass. CC can perform the rotation **file-to-file without printing the new value** (the `operator-present-one-time-secrets` convention) if asked — it did not do so unilaterally, because rotating a credential the operator holds in their own head or notes is their call, not CC's. **The reusable half, which matters more than this one password:** never use curl's `%{redirect_url}` (or `-v`, or `--libcurl`) against a basic-auth endpoint — all three re-render the credential. Confirm a redirect with `%{http_code}` and read the flash from a follow-up GET. | **Viktor decides**, CC executes |
|
||||||
| **R-399** | **How deep should the off-site integrity check go, and how often — VIKTOR RULES, CC EXECUTES.** The check ships at STRUCTURE depth (`monitoring.integrity.read_data_subset` empty). **⚠ THE FRAMING THIS ROW WAS FILED WITH WAS TOO NARROW AND THE MEASUREMENT SAYS SO.** It was scoped as a bandwidth-and-cadence question. It is more than that: **the structure check does not catch silent corruption at all.** Measured on demo-hp 2026-08-30 against a throwaway repo whose pack was corrupted WITHOUT changing its size — `restic check` returned `no errors were found`, **exit 0**; every `--read-data*` form returned `Pack ID does not match…` and exit 1. The structure check verifies the index, the pack inventory and the snapshot graph — it catches missing packs, broken indexes and unreadable snapshots, which are real failure modes — but it does **not** re-hash pack contents. **THE THREE NUMBERS, ALL MEASURED, NONE ESTIMATED:** live store **140 829 678 B (134.3 MB)**, 2 651 blobs, 67 snapshots; structure check **35.0 s**; and the full depth curve — 10% **35.9 s (+3%)**, 50% **37.3 s (+6%)**, 100% **39.2 s (+12%)**. **At today's size, re-reading ALL the data costs about four seconds more than reading none**, because the wall clock is dominated by SFTP round-trips over the WireGuard tunnel rather than transfer. **The caveat that keeps this honest:** these do NOT extrapolate — the structure check's cost tracks the INDEX, a read-data run's tracks the DATA, so a 50 GB store is ~370x the data and this curve says nothing about it. **WHAT HAPPENS IF YOU DO NOTHING:** the structure is checked weekly and **the data contents are never re-read**, so bit-rot inside a pack is not detected by anything in the product until a restore needs that pack. | **OPEN — DECISION (Viktor); CC executes in one config line** | — | Set `monitoring.integrity.read_data_subset` (accepted forms: `n/m`, `N%`, or a size like `50M`) and, if it should differ from weekly, `monitoring.integrity.max_age_days`. **Whoever turns read-data on must revisit `integrityCheckTimeout` (30 min)**, which was sized for a structure check and is noted as such at the constant. A larger store changes the arithmetic and this row should be re-measured before a fleet-wide default is chosen. | Viktor |
|
| **R-401** | **Revisit the off-site integrity depth when a real store is LARGE — the default rests on ONE measurement, on ONE 134 MB store.** Controller v0.228.0 (R-399) made `--read-data-subset=100%` the default for every box. The whole justification is a single data point: `demo-hp`, 2026-08-30, 140 829 678 B / 2 651 blobs / 67 snapshots, structure 35.0 s vs 100% **39.2 s** — four seconds. Re-proven live 2026-08-31 at 38.7 s. **It does not extrapolate, and the reason is structural: the structure check's cost tracks the INDEX, a read-data run's tracks the DATA.** A 50 GB store is ~370x the data and this curve says nothing about it. **Nothing was invented from that one point** — no rotation schedule, no size threshold, no bandwidth budget — because four production designs in this project were specced against unvalidated mechanisms and all four were wrong. **`readDataSubsetRe` already accepts `n/m`**, so a rotating schedule (`1/7` on a different seventh each week) needs no parser work when the time comes; the missing input is a measurement on a large store, not code. **THE TRIGGER IS AN EVENT, NOT A DATE:** the slow-check WARN from v0.228.0 firing on any box (`integritySlowNoticeThreshold`, 5 min) — that line names the duration, the depth and this row. **WHAT HAPPENS IF NOBODY ACTS:** every box re-reads its entire store every week, however large it grows, and the first person to notice is a customer whose upload is saturated. **Whoever acts must also revisit `integrityCheckTimeout` (30 min)**, which is now the number a large store meets first. | **OPEN — WATCHING** | — | When the WARN fires: measure the curve on that store, then choose between a rotation (`n/m`), a size-conditional default, or leaving it. Do NOT choose from this row's numbers — they are the small-store case. | CC |
|
||||||
| **R-400** | **The debug page has EIGHT buttons that post to endpoints which do not exist — not one.** Filed because v0.227.0 fixed one of them (`backup/integrity`, the seventh built-but-never-wired instance in this project) and the sweep the task asked for found the pattern is far wider than the single case. **Measured 2026-08-30** by comparing every `/api/debug/...` reference in `debug.html` against every `subpath ==` case in `handler_debug.go`: 24 referenced, 17 dispatched. The seven with no handler are **`backup/crossdrive`, `backup/infra`, `dr/infra-status`, `hub/infra-push`, `storage/simulate-disconnect`, `storage/simulate-reconnect`, `storage/watchdog-status`**. There is a single dispatcher, an exact-match `switch` with no prefix matching and a `default: http.NotFound`, so each of those buttons returns **404** — visible as an error rather than a silent success, which is the one mercy here. **`controller/README.md` documented four `backup/*` debug routes when only two existed**, corrected in v0.227.1. **Why this is a row and not a note:** a debug page is where an operator goes when something is already wrong, and a third of its controls do nothing. It is also the cheapest possible detector — a comparison of two lists — for a defect class this project has now hit seven times. | **OPEN — SMALL** | — | Per button: implement it, or delete it. **Do not leave the third state.** Then add the list-comparison as a gate — it is ten lines and it makes the eighth instance impossible on this surface. Note some may be deliberate stubs for features that moved to the agent (`backup/infra`, `dr/infra-status`); the answer per button is the finding, and deleting a button for a feature that lives elsewhere is still the right act. | CC |
|
| **R-402** | **The off-site integrity verdict and its depth are published to the hub and NO hub surface reads either.** `offsite.last_integrity_ok` has been on the wire since controller v0.227.0 and `offsite.last_integrity_depth` since v0.228.0; both are allowlisted in `scripts/wire_contract_gate.py` **with their reason**, which is why the gate is green rather than silent. **The order is deliberate and is the opposite of the one that produced R-331:** publish the value first, build the display when someone decides what the screen should say. R-331 removed a hub Backup card that rendered `Integrity Unknown` for every customer forever from fields nothing wrote. **The depth is not decoration:** "checked, OK" means two different things at structure depth and at 100%, so a card showing the verdict without the depth shows the same words for a check that re-read every byte and one that only read the index. **WHAT HAPPENS IF NOBODY ACTS:** the operator can only answer "was this customer's off-site store verified, and how deeply?" by reading that box's own log. | **OPEN — SMALL, needs a HUB decision first** | — | Decide what the hub screen should say, then model both fields hub-side and delete the two allowlist entries together. `offsite.last_integrity_check` is already decodable and is not allowlisted. | Viktor decides, CC builds |
|
||||||
| **R-362** | **A data drive detached mid-restore is reported as „permission denied".** Observed 2026-08-21 23:15: the guest-visible bind was unmounted 4 s into a scratch restore; the restore correctly failed and wrote nothing to the wrong place, but said „A visszaállítás sikertelen: restore dir: mkdir /mnt/felhom-drives/hdd_1/backups: permission denied". The controller has a drive-state concept (`IsDisconnected`, used by both backup legs) and the restore path never consults it. **A correct refusal that misdescribes why sends the reader at a permissions problem that does not exist.** Creditable in the same test: the agent re-bound the drive 5 s later, unaided. | **OPEN — MEDIUM** | — | Consult drive state when a restore path operation fails on ENOENT/EACCES and name the drive. | CC |
|
| **R-362** | **A data drive detached mid-restore is reported as „permission denied".** Observed 2026-08-21 23:15: the guest-visible bind was unmounted 4 s into a scratch restore; the restore correctly failed and wrote nothing to the wrong place, but said „A visszaállítás sikertelen: restore dir: mkdir /mnt/felhom-drives/hdd_1/backups: permission denied". The controller has a drive-state concept (`IsDisconnected`, used by both backup legs) and the restore path never consults it. **A correct refusal that misdescribes why sends the reader at a permissions problem that does not exist.** Creditable in the same test: the agent re-bound the drive 5 s later, unaided. | **OPEN — MEDIUM** | — | Consult drive state when a restore path operation fails on ENOENT/EACCES and name the drive. | CC |
|
||||||
| **R-363** | **The fill watcher runs once a day, so a filesystem that fills at 03:31 goes unannounced for ~24 h while the backup is already refusing apps.** `sched.Daily("fill-watch", "03:30", …)` (`cmd/controller/main.go:1092`) plus one startup check. Proven 2026-08-21 23:17: the 69 GB filesystem carrying the Docker data-root, the system namespace and ALL 40-class app data was filled to 99% / 1.2 GiB free; the backup reserve refused `kimai` per app and the hub received `recovery_unit_capture_failed` (error) naming the filesystem, **and the fill watcher said nothing at all**. The package comment says it "warns the CUSTOMER that a filesystem is filling, BEFORE anything fails"; at a daily cadence it frequently cannot. | **OPEN — MEDIUM** | — | The reserve already computes the same numbers every run. Let the watcher share that reading rather than owning a separate daily one. | CC |
|
| **R-363** | **The fill watcher runs once a day, so a filesystem that fills at 03:31 goes unannounced for ~24 h while the backup is already refusing apps.** `sched.Daily("fill-watch", "03:30", …)` (`cmd/controller/main.go:1092`) plus one startup check. Proven 2026-08-21 23:17: the 69 GB filesystem carrying the Docker data-root, the system namespace and ALL 40-class app data was filled to 99% / 1.2 GiB free; the backup reserve refused `kimai` per app and the hub received `recovery_unit_capture_failed` (error) naming the filesystem, **and the fill watcher said nothing at all**. The package comment says it "warns the CUSTOMER that a filesystem is filling, BEFORE anything fails"; at a daily cadence it frequently cannot. | **OPEN — MEDIUM** | — | The reserve already computes the same numbers every run. Let the watcher share that reading rather than owning a separate daily one. | CC |
|
||||||
| **R-364** | **Accented-text search is an instrument that silently transforms its input, and discipline alone has failed at least three times.** (1) 2026-07-20, `ssh → pct exec → bash -c`, nearly a wrong "banner cleared" claim (`felhom-controller/.claude/rules/ui-hungarian.md:19-22`). (2) 2026-08-13, `kubectl exec … sh -c grep` returned **0 for three strings that were present**, one step from a wrongly-reported failed hub deploy. (3) 2026-08-21, `tar -tf` rendered `őszibarack.md` as `\305\221szibarack.md`; recording the fixture's name bytes from that listing would have been wrong. **NOTE: that is two inside two weeks plus the founding case a month earlier — a third inside the two-week window is not on record.** | **OPEN — LOW** | — | **PROPOSED, NOT BUILT:** a helper that refuses to report a zero for any pattern containing a byte ≥ 0x80 unless a negative control also returns zero AND an ASCII anchor known to be present returns non-zero. Three probes, one helper, no judgement at the call site — because judgement is what failed. | CC |
|
| **R-364** | **Accented-text search is an instrument that silently transforms its input, and discipline alone has failed at least three times.** (1) 2026-07-20, `ssh → pct exec → bash -c`, nearly a wrong "banner cleared" claim (`felhom-controller/.claude/rules/ui-hungarian.md:19-22`). (2) 2026-08-13, `kubectl exec … sh -c grep` returned **0 for three strings that were present**, one step from a wrongly-reported failed hub deploy. (3) 2026-08-21, `tar -tf` rendered `őszibarack.md` as `\305\221szibarack.md`; recording the fixture's name bytes from that listing would have been wrong. **NOTE: that is two inside two weeks plus the founding case a month earlier — a third inside the two-week window is not on record.** | **OPEN — LOW** | — | **PROPOSED, NOT BUILT:** a helper that refuses to report a zero for any pattern containing a byte ≥ 0x80 unless a negative control also returns zero AND an ASCII anchor known to be present returns non-zero. Three probes, one helper, no judgement at the call site — because judgement is what failed. | CC |
|
||||||
|
|||||||
@@ -171,6 +171,15 @@ ALLOWLIST = {
|
|||||||
"that card. The sibling `offsite.last_integrity_check` is NOT allowlisted and passes on its "
|
"that card. The sibling `offsite.last_integrity_check` is NOT allowlisted and passes on its "
|
||||||
"own — the string already occurs hub-side. **When a hub surface is built, delete this entry.**"),
|
"own — the string already occurs hub-side. **When a hub surface is built, delete this entry.**"),
|
||||||
|
|
||||||
|
(_CH, "offsite.last_integrity_depth"): (
|
||||||
|
"R-399, controller v0.228.0: the DEPTH the verdict was reached at, published beside "
|
||||||
|
"last_integrity_ok above and unconsumed for the identical reason. It matters because 'checked, "
|
||||||
|
"OK' means two different things at structure depth and at 100%, so a hub surface that shows the "
|
||||||
|
"verdict without the depth shows the same words for a check that re-read every byte and one "
|
||||||
|
"that only read the index. Absent = the box cannot answer (a controller older than v0.228.0), "
|
||||||
|
"never 'structure'. **Delete this entry with its sibling when a hub surface is built** — filed "
|
||||||
|
"as R-402 in OPEN-ITEMS.md."),
|
||||||
|
|
||||||
(_AH, "wireguard.last_handshake_age_s"): (
|
(_AH, "wireguard.last_handshake_age_s"): (
|
||||||
"redundant: hub-side wgsync reconciles peers from its own state, and the OOB path's own "
|
"redundant: hub-side wgsync reconciles peers from its own state, and the OOB path's own "
|
||||||
"wg_handshake_age_s IS now decoded (into HostOOBRow, for the alert text)."),
|
"wg_handshake_age_s IS now decoded (into HostOOBRow, for the alert text)."),
|
||||||
|
|||||||
Reference in New Issue
Block a user