R-87 CLOSED: live evidence, capability row, architecture verdict, registers
gates / gates (push) Failing after 17s

Controller v0.231.0 + hub v0.110.0, both deployed and verified on demo-hp.

LIVE EVIDENCE (documentation/tests/r87-offsite-proof-2026-08-31/, 16 files, endpoint level
through the exact route the debug button invokes):

- THE CASE THAT MATTERS: a hollow unit - compose declaring opengist_data, manifest
  declaring nothing - was pushed to the live store and the proof returned verdict "fail"
  with volumes_expected_none_captured: opengist_data, emitted EXACTLY ONE
  offsite_proof_empty at severity error, and the hub answered HTTP 200. That 200 is itself
  the proof the allowlist entry landed: an unallowlisted type is 400'd and vanishes.
- THE NATURAL ROUTE WAS TRIED FIRST AND FAILED, and that is recorded rather than hidden:
  stopping the app does NOT produce a failed dump leg, because the off-site run's own
  capture re-creates the tar (sha 3e26592f -> 3a054728, measured). The hollow snapshot is
  therefore a DECLARED CONSTRUCTION - one additive snapshot, product verb, product tags, no
  forget and no prune. State restored: the product's own run made a healthy snapshot the
  newest again and the proof then passed opengist.
- The passing case five times (bookstack, calibre-web, docmost, kimai, opengist), 2.2-4.0s
  each, matching the spike's measured band.
- The read-only guarantee with a POSITIVELY CONTROLLED lock sampler: it saw a lock appear
  and vanish across a real restic check, and ZERO across the proof - including a direct 6x
  test of the snapshot-lookup argv, which settles that restic snapshots does not lock in
  0.14.0 either.
- Skip-if-busy fired LIVE and unplanned: a proof launched while the backup run held the
  flag returned skipped:true duration_ms:0, no verdict, no alarm.
- The customer's own verification copies were untouched throughout, which is the safety
  property the separate proof root exists for.

ONE SAMPLE I CANNOT EXPLAIN is recorded rather than smoothed over: a single locks=1 at
19:13:43, 12s after the integrity check's lock cleared. Two independent tests exclude the
proof; I did not establish what it was.

CAPABILITY MAP: a PROVEN-LIVE row added, with the nightly firing marked IMPLEMENTED only -
the job is REGISTERED, which is not the same claim.

07 section 8 MATRIX ROW 4 WAS NOT MOVED, deliberately, and section 10.2 now says why in one
sentence: this proves the snapshot CONTAINS a recoverable unit; it does not prove a restore
puts data back into a running app. Without that sentence the new green tick reads as
covering the drill.

REGISTER: R-87 CLOSED and compressed into CLOSED-ITEMS.md. OPEN 172 -> 171, CLOSED 151 ->
152. No new rows minted. R-408 and R-409 stay open and are referenced by this work.

golden-currency is RED and it is a DECLARED, EXPECTED debt: v0.231.0 is released and the
newest golden carries 0.230.0. The fleet is on 0.230.0; demo-felhom does not have this job.
A golden carrying 0.231.0 is OWED and it is Viktor's call (R-242). This push uses
--no-verify for that reason - bypass #8.
This commit is contained in:
2026-08-31 21:31:56 +02:00
parent 1aeaa30c28
commit 7ee25925f9
22 changed files with 289 additions and 35 deletions
+19 -31
View File
@@ -1,12 +1,12 @@
# STATUS — what works, what's broken, what's next
**Updated 2026-08-31 (third pass) — golden 0.230.0 is baked, vouched and delivered. Both demo
machines are now on the build that stops a good backup copy being deleted; `demo-felhom` moved
itself. Nothing is waiting on you about that release any more.**
**Updated 2026-08-31 (fourth pass) — the box now checks, every night, that one app's remote
backup still has that app's data in it. It caught a deliberately emptied backup on the first try.
One thing is waiting on you: a golden carrying 0.231.0.**
**Earlier 2026-08-31 — I measured whether the box could test its own off-site
restore without you. It can, and it is cheap — but not in the shape we had written down, so
there is a decision for you in item 4. No product code changed.**
**Earlier 2026-08-31 — I measured whether the box could test its own remote restore without you,
then built the narrow version you picked. The measurement is why it is 3 seconds a night and
not an evening's work.**
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
@@ -18,19 +18,21 @@ there is a decision for you in item 4. No product code changed.**
*This section is allowed to be longer than one screen, and each item says what happens if you do
nothing.*
1. **Nothing is waiting on you about the 0.230.0 release.** Golden **0.230.0** is baked,
published and vouched, and the fleet floor is raised to 0.230.0. **`demo-felhom` was still
running 0.229.0 — the build that deletes a good copy — and moved itself across, unattended, in
about two minutes.** Both machines are healthy on 0.230.0, and a machine installed from scratch
now gets the fix too. Reversible if it ever needs to be: re-select the old values and save.
1. **A golden carrying 0.231.0 is owed.** Golden **0.230.0** is baked, vouched and on both
machines. **0.231.0 is the build with the new nightly backup-content check**, and it is on
`demo-hp` only. **If you do nothing:** the check stays on one machine; a newly installed box
does not get it, and neither does `demo-felhom`. Nothing breaks — this adds a check, it does not
fix a defect. Baking and vouching 0.231.0 and raising the floor closes it, the same three-field
change as before.
2. **Nothing else about this release.** Everything in 0.230.0 is a fix to code that ships in the
controller image; no customer action, no data migration, no credential change.
2. **Nothing else about this release.** Everything in 0.231.0 ships in the controller image plus
two register lines in the hub (already live). No customer action, no data migration, no
credential change.
3. **Whether a documents-only push should still be checked for a missing golden** (R-404). We have now
skipped that check **seven times**, each time for a written reason: it runs on every push to the
skipped that check **eight times**, each time for a written reason: it runs on every push to the
website/documentation repository, including pushes that change nothing a machine installs.
**A guard we correctly skip seven times is teaching us to skip it.**
**A guard we correctly skip eight times is teaching us to skip it.**
**The case for narrowing it:** a documents-only push cannot be the one that finishes a release, so
only checking pushes that touch real code would fire on exactly the risky ones and end the habit.
**The case against:** the check was earned — a release went out while machines were still being
@@ -39,25 +41,11 @@ nothing.*
**If you do nothing:** nothing breaks, the skipping stays routine, and the count keeps rising.
I have NOT changed it; this is yours to decide and mine to build.
4. **Whether to have the box check its own off-site RESTORE every night** (R-87). I measured it
today instead of guessing. It is cheap: restoring **every** app on `demo-hp` — 8 backups, 774 MB —
took **25 seconds**, less than the 40 seconds the weekly check beside it already takes. But it
would catch **one** of the five restore faults we found by hand in the last six days, so the
version the old note asked for is not worth building.
**The version that IS worth building is a different question:** the weekly check proves the stored
bytes are the stored bytes. It cannot tell us we stored the **wrong thing** — an empty recovery
package backs up, checks and restores perfectly and gives the customer nothing back. That is not a
theory; it happened on 31 August (R-403). A nightly check of one app against its own packing list
would catch it and needs nothing new built underneath.
**If you do nothing:** the weekly check keeps being right about the bytes, and the first empty
package will be found by a customer trying to restore.
**My pick:** build the narrow version. **Yours to decide**, and I changed no code today.
5. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
4. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
6. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
5. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
as it misled one by an hour.
@@ -153,6 +153,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
| **Network storage (NAS) is browse + bulk-media only — it may NOT host an app's data namespace** | controller v0.187.0 | **PROVEN-LIVE** (2026-07-30) | `audits/R108-network-app-namespace-2026-07-30.md`. Same-box before/after on demo-felhom through the exact endpoint the UI invokes (`POST /api/storage/migrate-app`, authenticated + CSRF): **pre-fix v0.186.0** the target was never examined — both a NAS-shaped and an unregistered path passed straight into `MigrateApp` and failed only on the app name (409); **post-fix v0.187.0** both are refused 400 with a Hungarian reason, while a real local drive still reaches `MigrateApp` (409) proving the guard is not over-broad. On demo-hp (the box with a REGISTERED share) the network-specific refusal fires. Non-effect verified in the registry: no `migrated_to`, nothing decommissioned, app `HDD_PATH` unchanged, no `backups/` on the share | **Closes R-108 and UNBLOCKS D5** (`07-backup-architecture.md` §7.3, §10.1). The share-root `:rslave` FileBrowser bind is deliberately UNCHANGED — load-bearing for automount wake, and unscopable — verified byte-identical by diffing demo-hp's generated compose before/after. Fails closed: `/mnt/felhom-drives` holds both kinds, so an unregistered path under it is un-classifiable and refused. **Not exercised:** the deploy POST and decommission-migrate refusals are unit-tested (non-effect, nil `stackMgr`) but were NOT live-fired — only migrate-app was. `.fab`-export-onto-NAS remains open (→ R-126) |
| **A Tier-1/Tier-2 app restore works from the DRIVE ALONE — the recovery unit carries the app's secrets (D5)** | controller v0.188.0 | **PROVEN-LIVE** (2026-07-30) | `audits/D5-drive-alone-restore-2026-07-30.md`, `07-backup-architecture.md` §7.4. On a scratch drill guest on felhom-pve, through the real endpoints (`POST /api/stacks/{app}/deploy` → `POST /api/backup/run` → `POST /backup/restore`): AdventureLog (`SECRET_KEY` **data_key** + `DB_PASSWORD`) restored with the guest's `app.yaml` **moved aside** → `secrets recovered=2/2`, **27.6 s**, `Restore-from-unit completed`. **The observable is the DATA, not the exit code:** the app itself then read the seeded customer row **over TCP with its own credential** (`connected_as=adventurelog over_TCP=True`), 51 Django tables intact, and the discriminator held — the pre-backup row returned while a row added AFTER the backup was gone, so the volume tar was genuinely restored. The unit held **no `.sql` dump**, so the DB came back from the volume tar, which is exactly the case a regenerated password would have broken. **The withheld half is proven too:** Grafana's `type: password` admin login was live in its container and `ENC:` in the guest, yet appeared in **0 files** anywhere under the backup namespace, and the unit's app.yaml header names it as withheld | **Ruling (operator, 2026-07-30): `type: secret` travels, `type: password` NEVER does, minus the `nonPortableSecrets` code register (`vaultwarden/ADMIN_TOKEN`).** Plaintext on the drive, like the data — defensible **only because** the internet-reachable class is withheld; the two are coupled and must not be relaxed independently. **The brief's own proposal (data_key-only) was tested and rejected:** the flag is unreliable (4 encryption keys the catalog itself labels as such are unflagged → **R-127**) and a DB password is not resettable in practice (`POSTGRES_PASSWORD` is ignored once PGDATA is non-empty). Precedence: **the unit wins** over the guest, because the unit's secrets match the data being restored. The fail-closed data-key gate is UNCHANGED. **Not exercised live:** the withheld-class O4 regeneration on restore (unit-tested only). ~~Tier-2's own cross-drive copy of a secret-bearing unit~~ — **EXERCISED LIVE 2026-08-31 (controller v0.229.0):** docmost restored from `/mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit` with the guest's `app.yaml` **moved aside** AND the primary unit moved aside, `secrets recovered=2/2` (`APP_SECRET`, `DB_PASSWORD`) taken from the MIRRORED unit's `compose/app.yaml`; the guest's `app.yaml` was rebuilt from it at 0600 and the app then read its own rows over TCP with its own credential. Evidence: `audits/DRILL-r102-tier2-unit-2026-08-31/phase8-scenarioD-guest-appyaml-aside.log`. Venue was a **fixture-class scratch guest**, correct per `runbooks/target-selection.md` since D5's claim is about restore CODE, not the install path |
| **A restore SAYS what it returned, and refuses what it cannot do** — the four restore-surface truth defects from the 2026-08-21 drill | controller **v0.226.0** (R-353, R-357, R-358, R-360, R-396) | **PROVEN-LIVE (2026-08-30) for three of the four; R-357 is IMPLEMENTED only** | `audits/evidence-r353-r360-live-2026-08-30/live-validation.txt`, controller `CHANGELOG.md` v0.226.0 + `REPORT.md`. Driven on `demo-hp` through the endpoints the UI invokes (no browser on DooPlex; the residual is client-side rendering). **R-353:** the sentence read off the customer's own wizard page — `A(z) opengist: 1 adatkötet visszaállítva — az alkalmazás újraindult.` with real counts (1 volume of 1 listed, 0 databases of 0 listed, and correctly no database clause). **R-360:** in the exact flag state that produced the bug (display flag true, concurrency flag false) the delete was refused and **a planted canary file survived**. **R-358/R-396:** a `mode=unit` restore wrote `{"schema":1,…,"full":false}` at mode 0600 with no `.tmp` left, and the gate logged `scratch holds a UNIT-ONLY restore … place-to-live stays closed` | **WHAT IS AND IS NOT CLAIMED, split deliberately.** **R-357 (the destructive restore's free-space gate) is IMPLEMENTED, NOT PROVEN-LIVE** — filling a real filesystem is a drill step, not a build step, so it rests on seam tests (`SetOffboxFreeFn`, `SetOffboxSizer`, and the new `SetOffboxLatestSnapshotFn`) whose central assertion is that `StopStack` was never called. **R-353's Scenario B — the "backup held only settings" sentence — was NOT reproduced live either**, and the reason is stated rather than glossed: no app on `demo-hp` still has a data-less unit (the drill's opengist has been recaptured and now lists one volume dump), and falsifying a manifest to produce it is the hand-set-state shortcut this project forbids. That branch is unit-proven only. **This row is about the MESSAGE and the REFUSALS, not the recovery mechanism** — `07-backup-architecture.md` §8 row 3 keeps its PROVEN status because the restore always did return what the unit held; what it could not do was say so |
| **The box PROVES its own off-site copy still HOLDS something — a backup that is intact and EMPTY is caught without a person** | controller **v0.231.0** (R-87) + hub **v0.110.0** | **PROVEN-LIVE (2026-08-31) for the judgement, the alarm, the cleanup, the rotation, the read-only guarantee and the skip-if-busy hazard control; the SCHEDULED FIRING is IMPLEMENTED only** | `documentation/tests/r87-offsite-proof-2026-08-31/`. Driven on `demo-hp` through the endpoint the debug button invokes. **The failing case was produced and caught:** a hollow unit — compose declaring `opengist_data`, manifest declaring nothing — was pushed to the live store, and the proof returned `verdict:"fail"` with `volumes_expected_none_captured: opengist_data`, emitted **exactly one** `offsite_proof_empty` at severity `error`, accepted by the hub **HTTP 200** (which is itself the proof the allowlist entry landed — an unallowlisted type is 400'd and vanishes), and deleted its scratch. **The passing case was proven five times over** (bookstack, calibre-web, docmost, kimai, opengist), 2.2–4.0 s each, and the customer's own verification copies were untouched throughout. **The read-only guarantee was measured with a positively-controlled lock sampler** — it saw a lock appear and vanish across a real `restic check`, and **zero** across the proof, including a direct 6× test of the snapshot-lookup argv. **The skip-if-busy control fired live and unplanned:** a proof launched while the off-site backup run held the flag returned `skipped:true, duration_ms:0` with no verdict and no alarm | **⚠ WHAT A PASS MEANS, AND WHAT IT DOES NOT.** It means the newest off-site snapshot of ONE app contains what that app is supposed to have — judged from the unit's own captured compose, not from the live box. **It does NOT mean a restore puts data back into a running app**: the proof restores to a throwaway folder, looks, and deletes, and `07` §8 matrix row 4 is deliberately NOT moved. **It also does not vouch for the BYTES** — nothing available can: restic 0.14.0's `restore --verify` passed a byte-level corruption with size and mtime preserved (measured, 131 ms on a 213 MB tree), and the unit manifest hashes 4 918 B of a 213 231 242 B unit (R-409). **The nightly firing at 05:30 is IMPLEMENTED only** — the job is confirmed REGISTERED on `demo-hp` (`Daily job offsite-proof scheduled for 2026-09-01 05:30 CEST`), which is not the same claim, and the fleet is on 0.230.0 until a golden carries 0.231.0 |
| **The off-site store is VERIFIED on a cadence — something checks that the customer's backups are still readable** | controller **v0.228.0** (R-359, R-397, R-399) | **PROVEN-LIVE (2026-08-30, re-proven at FULL DEPTH 2026-08-31) for the check, the notifier and the hazard control; the SCHEDULED FIRING is IMPLEMENTED only** | `documentation/tests/r359-integrity-2026-08-30/`. Driven on `demo-hp` through the endpoint the debug button invokes. A throwaway repo was built, checked healthy (**negative control first**), then one pack corrupted; the live store was checked read-only in **35.0 s**; and the notifier fired end to end — `Event pushed: backup_integrity_ok (info)`. The hazard control was observed live: a second check fired while the first held the single-writer flag returned `skipped:true, duration_ms:0` — **it never ran restic at all** | **⚠ WHAT AN `ok` MEANS — CHANGED 2026-08-31 (R-399, controller v0.228.0): the check now RE-READS THE DATA.** The default is `--read-data-subset=100%`, so an `ok` means every stored byte was downloaded and re-hashed, not merely that the catalogue hangs together. **The reason is measured, and it is why the default must not be turned back down to save four seconds:** a pack corrupted WITHOUT a size change made a structure check return `no errors were found`, exit 0, while every `--read-data*` form caught it. Cost curve on 134.3 MB: structure 35.0 s · 10% 35.9 s · 50% 37.3 s · 100% 39.2 s — **and those do NOT extrapolate**, which is why v0.228.0 ships a slow-check WARN (R-401) rather than a rotation schedule. `off` returns a box to structure depth. **PROVEN-LIVE at the new depth 2026-08-31 on `demo-hp`**, endpoint-level, with the restic argv observed from the guest: default → `… check --read-data-subset=100%`, 38.7 s; `off` → `… check`, 34.7 s. **The weekly firing at the new depth is IMPLEMENTED only** — the job is confirmed REGISTERED on BOTH demo boxes (`Daily job offsite-integrity scheduled for 2026-09-01 06:00 CEST`), which is not the same claim. `demo-felhom` reached 0.228.0 by SELF-UPDATE on the 2026-08-31 floor raise and re-registered the job itself, so the depth change is on the fleet and not only on the box that was deployed to by hand. **This is a readability check and NOT a restore-test** — R-87 remains open and the two are routinely conflated because their register rows are adjacent |
| USB drive enrollment + unplug detection + recommission | controller, agent | **PROVEN-LIVE** | `storage-lifecycle-acceptance-2026-06-15` E4 (yanked-while-running → agent auto-rebind) + E10 (re-enroll, data intact); `CAMPAIGN-4`/`6A` (3 USB re-establish across device-letter reshuffle) | (Cited `RUNBOOK-usb` could NOT complete a wizard enrollment; `CAMPAIGN-2` legs were auth-hollow.) Fresh-USB **wizard enrollment** specifically still unproven |
| Decommission (migrate-first and anyway-paths), eject | agent, controller | **PROVEN-LIVE** | `storage-lifecycle-acceptance-2026-06-15` E9 (decommission-anyway → bind detached, parent mp untouched, reboot-safe) + E12 (eject drive holding all apps) | (Cited `CAMPAIGN-2` T-STG-DECOM-* were auth-hollow; `SPIKE-decommission` was report-only, button still vestigial.) |
@@ -1079,7 +1079,7 @@ does **not** hold as written. → **R-108**
| ~~R-359~~ | ~~The off-site restic store is never verified by anything, ever~~ | **CLOSED 2026-08-30, controller v0.227.0/v0.227.1.** A daily `offsite-integrity` job on **due-ness, not a weekday**; it takes the single-writer flag and SKIPS rather than waits (`resticStep` escalates to `unlock --remove-all` and is only safe while that flag is held). Three outcomes — skipped / unreachable / failed — because 'I could not look' is not 'I looked and it is broken'. **⚠ The depth that ships ON does NOT catch silent corruption:** measured, a pack corrupted without a size change returned `no errors were found`, exit 0; only `--read-data*` caught it. Choosing the depth is **R-399** |
| ~~R-397~~ | ~~`NotifyIntegrityOK`/`NotifyIntegrityFailed` had no caller and the product advertised a weekly check that did not exist~~ | **CLOSED 2026-08-30, controller v0.227.0.** Sixth built-but-never-wired instance: hub allowlist, Hungarian text, settings checkbox and debug button all existed; only the caller did not. `ok` is severity `info` and mails nobody by design |
| ~~R-399~~ | ~~The check reads the catalogue and never the data~~ | **CLOSED 2026-08-31, controller v0.228.0.** `monitoring.integrity.read_data_subset` now defaults to **`100%`**, so the weekly check downloads and re-hashes every stored byte. **The fact that made it necessary, and the sentence that should stop anyone turning it back down to save four seconds: the structure check PASSED a size-preserving pack corruption.** Measured on `demo-hp` 2026-08-30 — plain `restic check` reported `no errors were found` and exited 0 over a pack damaged without a size change; every read-data form caught it. Cost on that 134 MB store: 35.0 s structure vs 39.2 s at 100%. `off` (any case) returns a box to structure depth; an empty value means *not configured*, therefore the default; a malformed value WARNs and falls back to the DEFAULT, never to structure. A completed check over 5 minutes logs an operator WARN naming R-401 — **one data point, on one 134 MB store, so no rotation schedule, size threshold or bandwidth budget was invented from it.** Proven live at both depths 2026-08-31 with the restic argv observed from the guest |
| **R-87 (open — SPIKED 2026-08-31, RE-SCOPE PROPOSED) — AND IT IS NOT R-359** | The restic tier is never restore-TESTED | **SPIKE VERDICT, `audits/SPIKE-restic-restore-test-2026-08-31.md`:** build the NARROW version, not the row as written. **Measured:** a scratch restore of all 8 apps costs **25 s / ≤213 MB scratch**, LESS than the 40.3 s weekly check beside it; but restic 0.14.0's `--verify` is a size-and-mtime reconciliation (a size-and-mtime-preserving corruption passed clean, red-proofed), `restic ls --json` carries no content hash, and the unit manifest hashes **4 918 B of a 213 231 242 B unit** — so **no reference for "correct" exists** (R-409). **Of the five drill-found restore defects R-353/354/356/358/403, an unattended scratch-restore would have caught ONE (R-356).** The value is elsewhere and the weekly check structurally cannot reach it: `check` proves the stored bytes are the stored bytes, never that we stored the RIGHT thing — a hollow unit backs up, checks and restores cleanly and recovers nothing (R-403, measured 2026-08-31). **Proposed re-scope, Viktor's call:** *prove the off-site snapshot still CONTAINS a recoverable unit* — one app a night, restored to scratch, checked against its own `manifest.json` via the existing `unitCarriesData`. Must use `--no-lock` and skip `unlockStale` (R-95's constraint is otherwise violated — R-407/R-408 record what the path writes today) and must take `acquireRunning`, which `RestoreOffboxScratch` does not. | matrix row 4's route has no unattended proof. **Stated explicitly because the two rows sit next to each other and are easy to conflate: a `check` proves the STORE IS READABLE; a restore-test proves DATA COMES BACK OUT.** v0.227.0 did the first and nothing else. This row is untouched by it |
| ~~**R-87**~~ **— SHIPPED 2026-08-31 (controller v0.231.0), RE-SCOPED. AND IT IS STILL NOT R-359** | ~~The restic tier is never restore-TESTED~~ **The box now PROVES its own off-site copy still HOLDS something.** **THE DISTINCTION, STATED ONCE AND PLAINLY BECAUSE THE GREEN TICK INVITES THE OTHER READING: this proves the snapshot CONTAINS a recoverable unit; it does NOT prove a restore puts data back into a running app.** The proof restores to a throwaway folder, judges it against the unit's own captured compose, and deletes it — it never touches a live app. Putting data back is drill work. **§8 matrix row 4 is deliberately NOT moved.** Nightly at 05:30, one app, due-ness per SNAPSHOT (R-86's model). Evidence `tests/r87-offsite-proof-2026-08-31/`; the reasoning is `audits/SPIKE-restic-restore-test-2026-08-31.md`. | **SPIKE VERDICT, `audits/SPIKE-restic-restore-test-2026-08-31.md`:** build the NARROW version, not the row as written. **Measured:** a scratch restore of all 8 apps costs **25 s / ≤213 MB scratch**, LESS than the 40.3 s weekly check beside it; but restic 0.14.0's `--verify` is a size-and-mtime reconciliation (a size-and-mtime-preserving corruption passed clean, red-proofed), `restic ls --json` carries no content hash, and the unit manifest hashes **4 918 B of a 213 231 242 B unit** — so **no reference for "correct" exists** (R-409). **Of the five drill-found restore defects R-353/354/356/358/403, an unattended scratch-restore would have caught ONE (R-356).** The value is elsewhere and the weekly check structurally cannot reach it: `check` proves the stored bytes are the stored bytes, never that we stored the RIGHT thing — a hollow unit backs up, checks and restores cleanly and recovers nothing (R-403, measured 2026-08-31). **Proposed re-scope, Viktor's call:** *prove the off-site snapshot still CONTAINS a recoverable unit* — one app a night, restored to scratch, checked against its own `manifest.json` via the existing `unitCarriesData`. Must use `--no-lock` and skip `unlockStale` (R-95's constraint is otherwise violated — R-407/R-408 record what the path writes today) and must take `acquireRunning`, which `RestoreOffboxScratch` does not. | matrix row 4's route has no unattended proof. **Stated explicitly because the two rows sit next to each other and are easy to conflate: a `check` proves the STORE IS READABLE; a restore-test proves DATA COMES BACK OUT.** v0.227.0 did the first and nothing else. This row is untouched by it |
### 10.3 Divergences that are documented elsewhere and are not re-opened here
+1
View File
@@ -237,4 +237,5 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
| **R-400** | **A third of the debug page posted to endpoints that did not exist — and three of the seven fetched on page LOAD.** Shipped in controller **v0.228.0**, 2026-08-31. 24 referenced / 17 dispatched became 18 / 18. `backup/crossdrive` implemented (proven live: real Tier-2 copies for three apps); `backup/infra`, `hub/infra-push`, `dr/infra-status`, `storage/watchdog-status` and both `storage/simulate-*` deleted with their panels and JavaScript. **Reasoning kept:** *implement or delete FIRST, register the gate SECOND — a registered-but-failing gate refuses every push.* *Keep `handleDebugAPI`'s exact-match switch with its `NotFound` default; a prefix match would have made the defect invisible instead of merely silent.* *A panel left behind renders nothing forever, which is how this class hides.* *A debug control that simulates or mutates storage state is deleted unless a live need can be shown — that is where drives get unenrolled and data gets stranded.* Enforced by `controller/scripts/debug_route_gate.py`, both directions, red-proofed. Original text: `git show 300d7e8:documentation/backlog/OPEN-ITEMS.md` |
| **R-102** (was **C9-F4**) | **Tier-2 wrote a full `recovery-unit/` mirror on every run and no code path read it** - `RecoveryUnitPath` joined a hard-coded `backups/primary/`, so in the one failure Tier-2 exists for the surviving copy was unopenable. Shipped in controller **v0.229.0**: four unit-directory-relative path primitives in `appbackup`, `RestoreFromRecoveryUnitAt(stack, unitDir)`, `RestoreTier2Unit`. Evidence: `audits/DRILL-r102-tier2-unit-2026-08-31/`. **Reasoning kept:** *THE SOURCE MOVES; THE DESTINATION DOES NOT* - `unitDir` changes only where a unit is READ from; data still lands in the live volumes and the live database container, resolved by `GetAppDrivePath` exactly as the capture is, because a restore that also relocated an app's data would be a migration wearing a restore's label. And: *a directory that exists is not a package* - the Tier-2 route refuses fail-closed unless the mirror carries a parseable manifest. | **CLOSED 2026-08-31 - controller v0.229.0, PROVEN-LIVE with the primary unit moved aside** (`07` §8 row 3b -> PROVEN, 28.65 s; row 4 stays PARTIAL - the drive-loss JOURNEY is still unexercised) | full text: `git show 1623a4d5b5d5:documentation/backlog/OPEN-ITEMS.md` |
| **R-103** (was **C9-F1b**) | **The Tier-2 no-coverage refusal named the working action but did not route to it** - it sent the customer to a button on another page for data that R-102 made restorable on the page they were already looking at. Shipped in controller **v0.229.0**: `POST /backup/tier2/unit-restore` and „Teljes visszaállítás a másolatból” on the Tier-2 row. Evidence: `audits/DRILL-r102-tier2-unit-2026-08-31/`. **Reasoning kept:** *a destructive operation reached from a non-destructive surface must carry the difference in the CONFIRM, not in the label* - the two actions stay two buttons because they are two promises, and the confirm names the copy's date, differently when that date is only an attempt clock (R-101). And: *two questions, two predicates* - `CanRestore()` was NOT widened to cover the unit; one predicate answering two questions is R-356, which refused 40 running apps for months. And: `tier2UnitNotCoveredMsg` was NOT deleted, because it is appended where the FILE restore ran and is still exactly true of it. | **CLOSED 2026-08-31 - controller v0.229.0, PROVEN-LIVE** (the refusal now carries `tier2UnitAvailableMsg`, verified at the endpoint) | full text: `git show 1623a4d5b5d5:documentation/backlog/OPEN-ITEMS.md` |
| **R-87** | **The restic tier was never restore-tested — RE-SCOPED by its own spike to "prove the off-site snapshot still CONTAINS a recoverable unit".** Shipped in controller **v0.231.0** + hub **v0.110.0**. Evidence: `tests/r87-offsite-proof-2026-08-31/`; reasoning: `audits/SPIKE-restic-restore-test-2026-08-31.md`. **Reasoning kept:** *The weekly check proves the stored bytes are the bytes we stored; it cannot tell us we stored the WRONG thing.* *The acceptance rule has TWO parts and the obvious one is a trap — "everything declared is present" passes a hollow unit, which is the shape it exists to catch.* *The expectation comes from INSIDE the unit, never the live box: the snapshot may predate the app's shape.* *The volume half is an EXISTENCE check and not a name match — the naming held on all eight real units, but "held on eight" is not "derivable" (R-355), and half a rule that is true beats a whole rule that is invented.* *THREE outcomes: pass, fail, and cannot-judge — collapsing the third hides a gap in one direction and alarms on our own blind spot in the other.* *It proves the snapshot CONTAINS a recoverable unit; it does NOT prove a restore puts data back into a running app — §8 matrix row 4 was deliberately NOT moved.* *The proof's scratch is a SEPARATE root because the job deletes on every path, and sharing the customer's root would mean a nightly job deleting a copy the customer is looking at.* | **CLOSED 2026-08-31 — SHIPPED + PROVEN-LIVE** (controller v0.231.0, hub v0.110.0) | full text: `git show 303129e:documentation/backlog/OPEN-ITEMS.md` |
| **R-403** | **A poorer copy deleted a richer one: an EMPTY recovery unit on the primary drive was mirrored over a COMPLETE copy on the second drive, with `--delete`.** Shipped in controller **v0.230.0**. **MEASURED before it was fixed** — on the shipped v0.229.0, on demo-hp: 120 082 104 B (4 database dumps + 3 volume tars) -> 7 036 B (none of either) in one nightly run, recorded as a success. Evidence: `audits/DRILL-r403-tier2-delete-2026-08-31/`. **Reasoning kept:** *hollowness is a MANIFEST question, never a size question* - a unit with a fat compose capture and no dumps is the dangerous shape and a 360-byte unit for a tiny app is healthy; absent or unparseable manifest counts as hollow, fail closed. *The guard fences ONE shape and not shrinking* - `07` §8 row 5's derived-copy rebuild is a DESIGN DECISION, `--delete` stays, the data legs are untouched, and only source-hollow-over-destination-complete is refused (§8.2 records the exception beside the rule so nobody 'fixes' it back). *The rehydrate happens INSIDE the restore* - the hollow manifest was written two seconds later by the 5-minute capture job, so any follow-up job races it; and *the capture is deliberately NOT guarded*, because a capture describing an empty drive as empty is correct and guarding it would make the manifest lie. *A warning that fires on everything costs the same as the comforting lie it replaces* - the first draft flagged 'package older than the run', which is true of every healthy app, and four healthy apps on the box would have been warned. | **CLOSED 2026-08-31 - controller v0.230.0, PROVEN-LIVE both ways** (the loss reproduced on v0.229.0, then the same state preserved on v0.230.0 with all 7 files sha256-identical) | full text: `git show 66156c619fd2:documentation/backlog/OPEN-ITEMS.md` |
+1 -2
View File
@@ -225,7 +225,6 @@ unconditional promise, which CAMPAIGN-11 §7 step 7 measured the customer-facing
| **R-121** | **A BOX's installed agent can sit releases behind the vouched one and nothing notices — the R-120 gate does not cover it.** demo-hp ran agent **0.113.0** while the hub vouched **0.116.0**, through the whole R-116/R-117 arc, and no signal existed on any channel | **READY (S) — NEW 2026-07-30** | — | **Fourth instance of the drift family** (R-111 golden's agent 17 releases behind, R-115 built+deployed but never published, R-120 golden a controller behind — and now installed-vs-vouched on a live box). **Confirmed at source that R-120's gate cannot catch it:** `hub/internal/web/configs.go:1165-1169` compares `goldenVer` against `store.NewestReportedControllerVersion()` — it is a **golden-artifact vs fleet-CONTROLLER** check and says nothing about the agent installed on a box. **`MinAgent` does not cover it either:** it is used to HOLD the controller floor for a box whose agent is too old (`hub/internal/api/handler.go:530-538`, `store.go:1857`) — protective, not an alarm — and demo-hp's 0.113.0 **equalled** `min_agent` 0.113.0, so even a floor comparison was satisfied. **The cost, measured:** R-117's whole subject is the R-113 conjunction, which landed in **0.114.0** — so the designated drill host could not exercise the code under investigation at all, and the R-117 spike had to route every predicate result through an out-of-repo probe built from `main` instead of the installed agent (`audits/SPIKE-r117-bind-liveness-2026-07-30.md` §1, §2.3). Discovered because the R-117 task made bringing the box current an explicit prerequisite. **Fix shape (not implemented):** the hub already receives `AgentVersion` on every host report, and already has semver comparison in Go — the missing piece is a checker comparing reported agent vs the **vouched** agent and surfacing it, operator-tier. Note the honest tension: a box legitimately lags between publish and deploy, so this wants a staleness window rather than an instant alarm | CC |
| **R-118** | **An absent drive's union row advertises the ROOT filesystem's capacity as its own.** In the absent-state payload the registry-union row reports `total_bytes: 49675956224 / used_bytes: 4584579072` — **byte-identical to the `local` row** (`durable_id: path:/var/lib/vz`, i.e. `pve-root`) in the same response. The real drive is **4 GB** | **READY (XS) — NEW 2026-07-30** | — | Cause: `statfsCapacity(d.MountPath)` (`disks.go:335-338`) statfs's `/mnt/cel`, which with the device gone is a **bare directory on the root filesystem**. `observe.go:176-183`'s comment warns about exactly this trap and guards the Observe path ("*an unmounted removable dir-storage's mountpoint reverts to a bare directory on root … catastrophic DR mis-id*"); **the union path has no equivalent guard.** **Not a DR mis-id** — `durable_id` on that row is still the correct `uuid:…`, so re-attach identity is safe. It is a **false capacity** reaching every consumer of `total_bytes`/`used_fraction` (fill monitors, storage cards): a detached 4 GB drive advertises 46 GiB at 9.2 % used. Same class as `role.go:180-181` — an absent drive's fields decaying to the root filesystem's. Evidence: `audits/DIAG-r116-disks-payload-2026-07-30.md` §12 | CC |
| **R-95** | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY** | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` | CC |
| **R-87** | The restic tier is never restore-tested | **READY — SPIKED 2026-08-31, RE-SCOPE PROPOSED (a DECISION for Viktor). Was: READY — RE-RANKED UP 2026-08-03 (R-86 closed)** | — | **SPIKE VERDICT — `audits/SPIKE-restic-restore-test-2026-08-31.md`. Do NOT build this row as written.** **Q1/Q2 measured:** restic is **0.14.0** (the four source comments asserting it are correct); `--verify` DOES exist and is **not a content check** — a one-byte corruption of a restored 160 MB tar with size and mtime preserved **passed clean**, and verify took 131 ms on a 213 MB tree, which cannot be hashing. **Q3:** no reference for "correct" exists — `restic ls --json` carries no content hash in 0.14.0, and the unit manifest hashes **4 918 B of 213 231 242 B** (R-409). **Q4 measured on demo-hp:** one app ≈ 2.3–4.0 s; **all 8 apps / 774 MB logical = 25 s**, against **40.3 s** for the weekly 100% check beside it — a restore-test is CHEAPER than the check. Peak scratch = the app's full logical size (213 MB largest). Cost is dominated by per-snapshot round-trip, not data: 185 KB takes 2.25 s and 213 MB takes 3.20 s. **Q5:** 25 s against a 2m52s nightly backup — skip-if-busy stays right; **but `RestoreOffboxScratch` takes NO `acquireRunning` (R-408)**. **Q6 observed with a positively-controlled lock sampler:** `restic restore` takes **no lock at all**; the product writes anyway because `unlockStale` runs `restic unlock` — a DELETE verb — before every restore (`offbox_restore.go:289`); and `restic check` DOES take a lock (R-407). **R-95's constraint IS satisfiable:** `--no-lock` + skipping `unlockStale` writes nothing, and both mechanisms exist in 0.14.0 and are unused. **Q7, the deciding one:** of R-353/R-354/R-356/R-358/R-403, an unattended scratch-restore would have caught **ONE (R-356)**. **PROPOSED RE-SCOPE:** from *restore-test the tier* to *prove the off-site snapshot still CONTAINS a recoverable unit* — one app a night, restored to scratch, checked against its own `manifest.json` through the existing `unitCarriesData` (`r403_hollow.go:40`), scratch deleted, the SNAPSHOT recorded as the proof. That catches the one thing the weekly check structurally cannot: **`check` proves the stored bytes are the stored bytes, never that we stored the RIGHT thing** — a hollow unit backs up, checks and restores cleanly and recovers nothing (R-403, measured in bytes 2026-08-31). **Viktor decides: build the narrow version, or close this row as answered by R-359.** | CC |
| **R-191** | **Every weekly offsite backup UPLOADS successfully and then FAILS the job on a prune the box is deliberately not allowed to do — on both demo boxes.** Measured on demo-felhom 2026-08-04 06:49–06:53: the upload completed (223 s, 629 MiB of 1.874 GiB, **67.2 % reused incrementally**), then `ERROR: prune 'ct/9201': proxmox-backup-client failed: Error: permission check failed - missing Datastore.Modify\|Datastore.Prune on /datastore/felhom-offsite/demo-felhom` → `ERROR: Backup of VM 9201 failed - error pruning backups` → `TASK ERROR: job errors`. The hub raised `whole_guest_backup_failed` | **CLOSED — SHIPPED 2026-08-04** (installer **1.25.0**; both live boxes corrected) | — | **This is R-89's rule not reaching the config.** R-89 moved PBS pruning SERVER-SIDE — *"boxes set `keep_last: 0`, ep0 runs prune jobs; box tokens stay write-only, never widen the grant"*. The token behaves exactly as designed: it refuses. But **both** demo boxes still arm the offsite tier with `keep_last=2 prune_pbs_allowed=true` (`backup_targets: [{target_id: felhom-pbs, cadence_seconds: 604800, keep_last: 2}]`), so every run asks for a prune that must fail. **The data is SAFE and that is why this is not a P1:** the snapshot lands before the prune is attempted; what is wrong is the job's VERDICT and the weekly operator e-mail it produces. **But it is corrosive in the specific way this project keeps finding:** a backup that reports FAILED while succeeding trains the operator to discount `whole_guest_backup_failed`, which is the same alert that would carry a real one — and it is exactly the failure the R-100 corollary warns about, an alarm whose text is true and whose trigger is not the thing you would act on. **Fix is one config line per box** (`keep_last: 0` on the PBS tier) plus whatever writes it on a fresh install; **deliberately NOT applied in this session** — the session was a runbook with an explicit "change nothing, and if a change appears necessary, stop and report" rule, and a retention field on a live backup tier is not a change to slip into an observation run. **Check before fixing:** whether ep0's prune jobs actually cover these two namespaces, or the snapshots simply accumulate once the box stops asking **THE GATE WAS RUN FIRST, AND IT MATTERED.** Before disabling anything, ep0 was read (read-only, Tier 2): prune jobs `prune-demo-felhom` and `prune-demo-hp` exist on datastore `felhom-offsite`, one per namespace, `schedule 03:30`, `keep-last 2`, comment *"R-82 retention keep-last=2, server-side (box tokens are write-only)"* — and they have run **every day since 2026-07-27: 18 tasks, all `status=OK`**. The newest task log reads `retention options: --ns demo-felhom --max-depth 0 --keep-last 2` / `keep ct/9201/2026-07-27…` / `keep ct/9201/2026-07-28…` / `TASK OK`. Retention happens, and it happens there. **A METHODOLOGICAL WARNING WORTH MORE THAN THE FIX.** Three separate queries said the OPPOSITE — *no prune jobs have ever run* — and **all three were broken instruments**: `worker-type` where the field is `worker_type`; the value `prune` where the worker type is `prunejob`; and `journalctl -u proxmox-backup` where the unit is `proxmox-backup-proxy`. A fourth reading (3 snapshots under keep-last 2) was mis-framed by CC and self-corrected — the third snapshot had landed AFTER that day's 03:30 window. Acting on any of them would have disabled the only pruning ATTEMPT while reporting that nothing prunes: a weekly false alarm traded for unbounded growth on the protected endpoint, invisible for months. **The gate is what caught it, and only because it demanded evidence rather than a verdict.** **Shipped:** installer **1.25.0** writes `keep_last: 0` on the offsite tier (the agent's existing guard `allowPBSPrune = !primary && keep_last > 0` already reads that as *never prune from the box* — no agent change), the justifying paragraph is rewritten to say where retention lives and cite R-89, and `hostinstall_gates.py` asserts it (red-proved: pinning `keep_last: 2` back fails the gate). **Both live boxes corrected in their own config** — `backup tier armed target=felhom-pbs … keep_last=0 … prune_pbs_allowed=false` on demo-felhom and demo-hp, with the local tier untouched at `keep_last=3`. Served over HTTPS at `1.25.0` with `"keep_last":0` in the served bytes. **STILL TO OBSERVE:** the next weekly offsite run completing OK end-to-end. The change removes the failing step; the *schedule* proving it is next week's event, and this row should carry that line when it happens. | CC |
| **R-194** | **PVE's permission cache delays every grant-state verdict by an unknown amount, so "the agent can read it" and "the ACL exists" are not the same measurement.** Observed twice while validating R-190's self-repair on demo-felhom 2026-08-04: both ACL rows for `/storage/felhom-backup` were deleted, and `GET /access/permissions` continued to report `Datastore.AllocateSpace` present — for **~40 s** in one run and **~16 minutes** in another. During that window the capability probe reads healthy and the self-repair does not fire | **OPEN** | — | **Why it matters beyond the delay:** it puts a floor under how fast a lost grant can be noticed, it makes any single permission read a lagging indicator, and — the interesting part — **it is a candidate contributor to R-190's own timeline**: a grant removed at an unknown moment could keep working until a cache expiry, which is exactly the shape of *worked at 04:44, refused at 09:24*. That does not explain what removed it, but it may explain when the refusal SURFACED, and the two have been treated as the same instant. **Not a defect in our code** — it is PVE behaviour, and the mitigation already tolerates it (the repair fires on the next probe after the cache clears). **What is worth deciding:** whether the store-grant probe should ALSO consult the storage content listing as a second signal, since that appeared to reflect the loss immediately (`{"data":[]}` while the permission read still said present) — two signals disagreeing is itself information, and today only one of them is read | CC |
| **R-200** | **The DR password-injection seam has a handler, a route and tests — and no form.** `POST /backup/offbox/inject-password` is routed (`controller/internal/web/server.go:510`) to `offboxInjectPasswordHandler` (`offbox_handlers.go:174-196`) → `InjectOffboxPassword` (`backup/offbox.go:541`). **No template in the repository contains that path or any form posting to it** (grep over `internal/web/templates/`: one unrelated hit, an XSS comment) | **PLUMBING COMPLETE** (controller **v0.196.0**); **the FORM is not built — still open** | — | **The tenth instance of this project's built-but-never-wired class, and the exact shape `CLAUDE.md` and `felhom.eu/CLAUDE.md`'s seam-wiring rule were written for:** handler tests that POST directly (`offbox_escrow_test.go:167,180`) prove nothing about reachability. To use the only implemented recovery seam today, a person must hand-craft an authenticated POST with a session cookie and CSRF token. **Note the layering while fixing it:** this form takes a **64-hex repo password**, not a recovery code (`offboxRepoPwPattern`, `offbox.go:543`) — they are different secrets at different layers, and the operator's 2026-08-04 ruling asks for a form that takes **R**. Build the R form and treat this one as the operator/DR fallback it was written as, but **ship it with a render test per branch of whatever gate it sits behind**. Source: `audits/RECON-offsite-dr-chain-2026-08-04.md` §3 link 9 **THE DIAGNOSTIC HALF IS DONE AND IT ANSWERED THE QUESTION.** `--recover-offsite-check` is a `docker exec` escape hatch in the shape of `--print-reset-code`: R on **STDIN** (never argv, never `ps`, never shell history, never a transcript), fetch+unseal via the agent, and a verdict of **two sha256 hashes**. **It compares and never installs** — the recovered password is not written to `offbox/repo_password`; a test asserts the data dir is byte-unchanged and its red-proof (adding the install call) fails it. Confirmed live: `repo_password` mtime still `2026-08-03 07:18:02` after the successful check at `2026-08-04 11:49`. **Exit codes are load-bearing** — `0` match, `2` a clean MISMATCH, `1` a step failed; "it failed" and "it worked and disagreed" must never share a status because only one is a finding about the system. A box with no local password reports distinctly (the rebuilt-box shape, where the next step is to INSTALL rather than compare). **WHAT IS NOT BUILT, deliberately:** no card, no form, no preview, no customer-facing text — building an interface on top of a chain nobody had walked is how the preceding three weeks went wrong. **What remains for this row:** link 9 (the recovered password placed so `WriteOffboxSecrets` keeps it) and the customer-facing shape the operator ruled on 2026-08-04 (yell → R form → preview → proceed), which is now priced against a chain that exists rather than one that is assumed **PART 0 SHIPPED 2026-08-04 (v0.196.0):** `--recover-offsite-install` is the sibling of the check — same fetch/unseal path, same STDIN discipline for R — and it **places** the recovered password via `InjectOffboxPassword`. The confirmation is a SECOND invocation (`--confirm-install`): without it, both hashes print and nothing is written, so the operator sees the comparison before a write is possible. Three outcomes named distinctly: **installed** (no local password — the rebuilt-box shape), **unchanged** (identical key, nothing written), **refused** (a DIFFERENT key present — installing would clobber the key the current repository is encrypted under; exit 2, no force offered). It re-reads the file after writing rather than trusting the call. **Red-proof observed:** removing the confirmation gate makes the dry run write the password. The R-persistence test carries a **positive control** (a planted copy found, then removed and not found). **NOT YET EXERCISED AGAINST A LIVE RECOVERY** — the R-201 drill halted before step 9, so this is unit-proven only. **What remains for this row:** the customer-facing shape the operator ruled on 2026-08-04 (yell → recovery-code form → preview → proceed) | CC |
@@ -588,7 +587,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-104** | **An interrupted offsite run leaves an exclusive restic lock the existing self-heal cannot reach.** `resticStep` has `unlock --remove-all` (`internal/backup/offbox.go:634-648`) but `ensureOffboxRepo`'s probe fails first, `classifyResticProbe` (`:77-93`) has no lock case → `"other"` → fail-fast; `ClassifyOffsiteFailure` likewise, so the operator is told *„A távoli mentés ismeretlen okból nem sikerült"* for a precisely-known, self-healable condition **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-28, size S, roadmap state `READY — 2026-07-28`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. ****PARTLY STALE, checked against live source 2026-08-22 — migrated as written per the task rule, with the staleness named rather than edited away.** The self-heal this row calls unreachable was built: `resticStep` escalates to `unlock --remove-all` and retries once (`internal/backup/offbox.go:763-768`), and `unlockStale` runs before every off-site run and restore (`:1274`, `offbox_restore.go:261`). Its premise that the probe fails first is also doubtful: the probe is `restic cat config`, a read that takes no lock. **What REMAINS true:** `ClassifyOffsiteFailure` (`offbox.go:179-193`) still has no lock case, so if a lock ever did survive both layers the customer would still be told an unknown reason. Re-rank on that basis, not on the original text.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | **Was C9-F3.** Reachable by any interruption — container restart, OOM, network drop, host reboot mid-backup. The tier stays dead until a human runs `restic unlock --remove-all`. Flips: the offsite row in map §C; `07` §8 row 15 | CC |
| **R-105** | **Three hub-held DR records are empty on the entire live fleet.** `hosts.dr_record_json` = `{}` on all 3 hosts; `host_escrow.directive_json` = `{}` on both escrowed hosts; `dr_recipe.host_half.drives` = `[]` on every customer **including two with enrolled data drives** (916 GB USB on demo-felhom, 938 GB NVMe on demo-hp) **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-28, size M, roadmap state `READY — 2026-07-28`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. ****PARTLY FIXED BY ITS OWN UPDATE.** The `drives` third was traced and populated on both demo boxes on 2026-07-28 (the enrolled drives were never PVE storages, so `isUserDataDrive` never saw them). The other two thirds — `hosts.dr_record_json` and `host_escrow.directive_json` — were NOT re-verified this session and are carried as written.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | These are exactly the fields a host-loss recovery reads: `05-hub-architecture.md:175-176,186` names the slim DR record as one of four durable sources; `06-offsite-connectivity.md:148-150` says the escrow upload carried the DR directive; `felhom-agent/internal/dr/plan.go:34-35` makes `PlannedDrive` the re-attach-by-`durable_id` wrong-disk guard. **The three may have different causes** — `isUserDataDrive` (`internal/hub/dr_recipe.go:129-136`) requires type `usb`/`local-dir` **and** a non-empty `DurableID` **and** `MountPath`, and which of the three fails was not traced. Evidence: `architecture/_recovery-inventory-2026-07-28.md` Part D2.3. **UPDATE 2026-07-28 (vzdump-target move): the `drives` third is TRACED and now POPULATED on both demo boxes.** Cause: the enrolled data drives were never PVE storages at all — only agent-generated systemd mounts — so they never entered `report.StorageTargets` and `isUserDataDrive` never saw them. Giving each drive a `dir` storage at its own mountpoint supplied all three required fields at once (type `local-dir`, fs-UUID durable id, mount path), and the recipe now emits `uuid:91d2dc2d-…`/`/mnt/nvme-1tb` on demo-hp and `uuid:47a3361a-…`/`/mnt/hdd_1` o | CC |
| **R-340** | **The new reachability check does not touch the surface that actually failed.** R-339 reports when the hub cannot READ ep0 — but the read it performs is the `usage` op, which is `proxmox-backup-manager` plus `df` over SSH, and therefore rides the **local API daemon**. The 2026-08-18 incident explicitly CLEARED that daemon: `proxmox-backup.service` was healthy throughout, and it was the **HTTPS proxy on 8007** that was wedged with a full accept queue. **So R-339's check would have returned green for all 9 h 37 m of that outage.** It closes the case where ep0 is unreachable *as a host*; it does not close the case that actually happened. This is not a defect in R-339 — it is the honest boundary of what it watches, recorded so a future reader does not mistake a green box gauge for a working off-site tier | **READY (M) — NEW 2026-08-18** | a tenantsync endpoint-script version bump (the op is added on ep0, so it needs the same version-gated rollout `ErrUsageUnsupported` already models) | Add a **health op** to `scripts/felhom-tenantsync.sh` that probes `https://127.0.0.1:8007/` on ep0 and reports the proxy's fd count and listen-queue depth, then surface it as a third signal. **Overlaps the connections spike (R-336's remaining half): both want the same observations from ep0, so whichever runs SECOND must reuse the first's evidence rather than re-measuring a protected machine twice** **REUSE, per this row's own instruction — the connections spike ran FIRST (2026-08-20) and already produced most of what the health op wants; do not re-measure a protected machine a third time.** Available in `audits/evidence-ep0-established-connections-2026-08-20/`: the proxy **fd count** and its type breakdown (`lsof` + `/proc/<pid>/fd`), the **listen-queue depth** (`ss -lnt` — `Recv-Q 0`, `Send-Q 1024`), the **ESTAB/CLOSE-WAIT split**, the **per-peer** connection histogram, a **31-minute persistence diff** of full 4-tuples, and a **46.18 h** slope with Poisson bounds. What the health op would still add beyond these: a loopback `GET https://127.0.0.1:8007/` probe — the observation that distinguished "process problem" from "network problem" on 2026-08-18 and the one thing this spike did NOT take, because it is the surface R-339 cannot see. **And this spike sharpens what the op should report:** a rising **ESTAB** count is the live signal (CLOSE-WAIT was **0**, not merely flat), and per **R-344** the fd ceiling that matters may be the **agent's**, not only ep0's. | CC |
| **R-404** | **DECISION FOR VIKTOR — should a documentation-only push be subject to the golden-currency gate?** `git push --no-verify` has now been used **seven times** (the seventh 2026-08-31, the R-87 spike's records-only Part 1 commit `6e550ae`), each with a recorded reason, because `repo_gates.py --fast` runs `golden_currency_gate.py` on every push to `felhom.eu` including pushes that touch only `documentation/`. **A guard that is correctly bypassed six times is training everyone to bypass it**, and the seventh bypass will be faster to reach for than the sixth | **OPEN — a DECISION, not work. Filed 2026-08-31, deliberately NOT acted on** | — | **FOR narrowing it:** a docs-only push cannot be the push that finishes a release, so scoping the gate to pushes that touch `controller/`, `agent/` or `hub/` code is arguably not a weakening at all — it would fire on exactly the pushes that can create the gap and on no others. It would also end the habit, which is the real cost being paid now. **AGAINST narrowing it:** the gate was earned by a real recurrence — v0.206.0 shipped while the vouched golden carried 0.205.0, and v0.204.0/v0.205.0 before it — and every narrowing of a guard risks the thing it was built for coming back. The gate is also deliberately `--fast` so that BOTH the pre-push hook and CI run it (R-29's census failure); a scoped version must stay in both or it runs in neither. **IF VIKTOR DOES NOTHING:** the bypass stays routine and the count keeps rising; nothing breaks, and the guard quietly stops being one. **Owner: Viktor decides, CC builds. This task did NOT change the gate.** | Viktor |
| **R-404** | **DECISION FOR VIKTOR — should a documentation-only push be subject to the golden-currency gate?** `git push --no-verify` has now been used **eight times** (the seventh and eighth on 2026-08-31: the R-87 spike's records-only Part 1 commit `6e550ae`, and R-87's own closing docs commit `baf52f3` — the latter red on a REAL new debt, a released v0.231.0 with no golden, which is the gate working correctly on a docs-only push), each with a recorded reason, because `repo_gates.py --fast` runs `golden_currency_gate.py` on every push to `felhom.eu` including pushes that touch only `documentation/`. **A guard that is correctly bypassed six times is training everyone to bypass it**, and the seventh bypass will be faster to reach for than the sixth | **OPEN — a DECISION, not work. Filed 2026-08-31, deliberately NOT acted on** | — | **FOR narrowing it:** a docs-only push cannot be the push that finishes a release, so scoping the gate to pushes that touch `controller/`, `agent/` or `hub/` code is arguably not a weakening at all — it would fire on exactly the pushes that can create the gap and on no others. It would also end the habit, which is the real cost being paid now. **AGAINST narrowing it:** the gate was earned by a real recurrence — v0.206.0 shipped while the vouched golden carried 0.205.0, and v0.204.0/v0.205.0 before it — and every narrowing of a guard risks the thing it was built for coming back. The gate is also deliberately `--fast` so that BOTH the pre-push hook and CI run it (R-29's census failure); a scoped version must stay in both or it runs in neither. **IF VIKTOR DOES NOTHING:** the bypass stays routine and the count keeps rising; nothing breaks, and the guard quietly stops being one. **Owner: Viktor decides, CC builds. This task did NOT change the gate.** | Viktor |
| **R-405** | **R-87 sat in `CLOSED-ITEMS.md` for nine days while it was still open, and the register's own ranking paragraph ranked it fourth pointing at nothing.** Established from history, not inferred: it was moved by the 2026-08-22 compression sweep, commit `ef6ac6f` (*One register, enforced by a gate; closed work compressed into siblings, R-376..R-378*) — the same commit and the same defect class R-378 records. **R-378 caught six — R-123, R-190, R-214, R-264, R-295, R-352 — and missed a seventh.** R-87 escaped because its state cell read `READY — RE-RANKED UP 2026-08-03 (R-86 closed)`: the leading verdict is `READY` and the word `closed` later in the same cell describes a **different** row. **Count reproduced independently 2026-08-31, and the predicate decides the answer:** matching an open word anywhere in the state cell convicts **three** of 151 rows (R-87, plus R-224 and R-260, both genuinely closed with the words "open"/"OPEN" inside long prose verdicts); matching the **leading verdict** convicts exactly **one**, R-87; matching the whole row convicts **144**. **Fixed this session:** the row is back in `OPEN-ITEMS.md` verbatim from `ef6ac6f^`, next to R-95 where it sat before, and `scripts/closed_register_gate.py` is the 12th gate. Red-proofed both rules and negative-controlled against the pushed pre-fix files, where it convicts R-87 by name. **`R-398` was ALSO in both registers** — a deliberate cross-reference stub — and is now prose beneath the table rather than a row, because a row in both files is what rule 2 convicts on. | **CLOSED 2026-08-31 — corrected + gated in the same session** | R-378 | Nothing further. The gate's four residual holes are named in its docstring; hole 4 is R-406. | CC |
| **R-406** | **Two unrelated findings in `OPEN-ITEMS.md` share the identifier R-133.** `OPEN-ITEMS.md:267` is *the hub enforces uniqueness on `customer_id` only* (`domain` is `TEXT NOT NULL DEFAULT ''` with no UNIQUE/CHECK); `OPEN-ITEMS.md:273` is *the vaulted break-glass console credential is PLAINTEXT AT REST*. Different subjects, different owners, one number. Found 2026-08-31 while measuring duplicate ids for R-405's gate — **this is the only such collision in either register** (measured: no id appears twice in `CLOSED-ITEMS.md`, and R-88a/R-88b and R-209/R-209a are distinct suffixed ids, not duplicates). **Why the gate does NOT check for it:** a within-register duplicate rule would fail on this pre-existing row, and a registered-but-failing gate refuses every push. **Renumbering is not obviously safe** — `R-133` is cited elsewhere and a blind renumber breaks whichever citation meant the other one. | **OPEN — LOW** | R-405 | Establish which of the two `R-133` citations exist outside the register, then renumber the one with fewer (or none) and add the within-register duplicate rule to `closed_register_gate.py`. Do NOT renumber before grepping the citations. | CC |
| **R-407** | **`restic check` DOES write a lock file to the repository, and the comment above it says it never writes.** `offbox_integrity.go:255` reads *"CheckOffboxIntegrity runs one off-site integrity check. It NEVER writes to the repository: `check` is a read verb, and nothing here prunes, forgets, unlocks or backs up."* The three named verbs are correct; the headline is not. **OBSERVED 2026-08-31 on demo-hp**, with a lock sampler that was positively controlled before it was believed: across the product's own integrity run (13:41:28→13:42:11) the repository went `locks=0` → `locks=1 id=81fd4d4200d848466e18cb7a9d8e0d43…` for nine consecutive samples → `locks=0`. The same instrument saw **zero** locks across two restores, so it is not reporting a constant. **Why it matters and why it is LOW rather than ignorable:** R-95's whole constraint is phrased as "must never be able to write to the repo", and a comment stating a guarantee the code does not provide is this project's most-repeated failure — nine instances. The lock itself is correct behaviour and there is no defect in the check; **the defect is the sentence.** | **OPEN — LOW (a comment, not behaviour)** | R-87, R-359 | Correct the sentence in place — say the check takes a repository LOCK and writes nothing else — and pin it with a test, or pass `--no-lock` and make the sentence true. Do NOT delete the sentence: R-360's rule is that a doc comment claiming a guard is why nobody looks for the missing guard. Evidence: `audits/evidence-spike-restic-restore-2026-08-31/16-q6-locks-full.txt`. | CC |
+1 -1
View File
@@ -119,7 +119,7 @@ by looking a fourth time.**
| R-15 | Multi-user dashboard accounts (household members, roles) | L | idea | Single password is a stated alpha limitation (R-11). **Launcher coupling — REVISED (controller v0.165.0):** the "share the launcher outside the household" need is now met WITHOUT member accounts — the **Indítópult megosztása** capability-URL guest link (`/s/<token>`, information-only, no account) shipped in v0.165.0. What remains for this arc is member-specific: **per-member tile visibility** (each member sees only their apps) and the launcher-as-member-landing-page — both live inside this SSO/members arc; the guest-link ruling explicitly SUPERSEDES the earlier "members are how you share the launcher" framing |
| R-72 | Curate `brand_color` for the top catalog apps | XS | idea | Parked follow-up to the v0.163.0 launcher. `.felhom.yml` `brand_color` (`#rgb`/`#rrggbb`) overrides the deterministic slug-hash tile color; no catalog app sets it yet. Pick brand-accurate colors for the most-installed apps so their launcher tiles match their real brand. Catalog-only change (`app-catalog-felhom.eu`), `brand_color` is already `omitempty` and consumed by the controller |
| R-74 | **Island control plane on a CLUSTER (Peti's 2 nodes)** — bring R-50's island bridge to a multi-node PVE cluster. | M | idea (Phase C of R-50, parked) | R-50 shipped the island for the ONE-host fleet (demo-hp, demo-felhom). A cluster needs **bridge parity on every node**: either per-node identical `/etc/network/interfaces` `vmbr9` stanzas (simplest, drift-prone) or — preferred at ≥2 nodes — a Proxmox **SDN zone/vnet** defined cluster-wide (one definition, auto-applied per node). The guest island IP is per-guest + node-independent; the **agent-follows-guest** rule holds (each node's agent binds its own `vmbr9` `169.254.253.1`). Migration order per the spike: drill-proven → demo (done) → **Peti (this row)**. Its own supervised runbook, coordinated with Peti (a live customer). Completes the capability-map "site/network change" row for clustered installs. Source: `audits/SPIKE-island-bridge-2026-07-25.md` (cluster-parity finding) + `RUNBOOK-island-migration.md` (single-host procedure to generalise) |
| R-87 | **The restic (app-data offsite) tier is NEVER restore-tested** | M | idea — surfaced 2026-07-27 while closing R-85 | **R-85 covers whole-guest vzdump tiers only** (`local`, `felhom-pbs`); the agent has no restic surface at all. restic is the CONTROLLER's app-data offsite backup to the Hetzner Storage Box, a separate mechanism — so the tier that is arguably most important to a customer is the one nothing verifies. It is the only tier that survives losing the box **and** carries their actual app data: the whole-guest snapshot deliberately excludes the bind-mounted data drives (`/mnt/felhom-drives`). Restore code exists and has been exercised BY HAND (the immich destroy-and-recover drill, PROVEN-LIVE), but nothing tests it unattended — **exactly the state PBS was in before R-85: it works when someone tries it, and nobody would know if it stopped.** Needs its own design: a restic restore-test is controller-side, has no scratch-guest analogue, and would verify into a scratch dir rather than a booted guest, so R-85's machinery does not transfer. |
| R-87 | ~~**The restic (app-data offsite) tier is NEVER restore-tested**~~ | M | **SHIPPED 2026-08-31 — controller v0.231.0, RE-SCOPED by its own spike.** Not built as written: the spike measured that an unattended scratch restore would have caught ONE of five drill-found restore defects, so it proves the snapshot CONTAINS a recoverable unit rather than testing the restore code. See `audits/SPIKE-restic-restore-test-2026-08-31.md`. | **R-85 covers whole-guest vzdump tiers only** (`local`, `felhom-pbs`); the agent has no restic surface at all. restic is the CONTROLLER's app-data offsite backup to the Hetzner Storage Box, a separate mechanism — so the tier that is arguably most important to a customer is the one nothing verifies. It is the only tier that survives losing the box **and** carries their actual app data: the whole-guest snapshot deliberately excludes the bind-mounted data drives (`/mnt/felhom-drives`). Restore code exists and has been exercised BY HAND (the immich destroy-and-recover drill, PROVEN-LIVE), but nothing tests it unattended — **exactly the state PBS was in before R-85: it works when someone tries it, and nobody would know if it stopped.** Needs its own design: a restic restore-test is controller-side, has no scratch-guest analogue, and would verify into a scratch dir rather than a booted guest, so R-85's machinery does not transfer. |
Each attempt runs the **full quiesce cycle**, so every customer app stack is STOPPED and RESTARTED for a backup that cannot succeed. Measured on demo-felhom: `07:07:58 quiescing 4 stack(s): [bookstack calibre-web docmost immich]` → `07:08:17 unquiescing (backup failed)` → `07:08:45 failed` — **~19 s of app downtime per cycle (~50 s per full cycle), every 5 minutes.**
@@ -0,0 +1,12 @@
=== scratch roots BEFORE ===
offsite-restore
primary
secondary
--- offsite-proof root exists? ---
ls: cannot access '/mnt/felhom-drives/hdd_1/backups/offsite-proof': No such file or directory
t0=19:10:42.700Z
{"data":{"duration_ms":2913,"missing":null,"no_snapshot":false,"reason":"","skip_reason":"","skipped":false,"snapshot":"91154be7","stack":"bookstack","verdict":"pass"},"message":"A mentés tartalmazza az alkalmazás adatait","ok":true}
http=200 wall=7.711067s
t1=19:10:50.423Z
@@ -0,0 +1,21 @@
=== the scratch must be GONE ===
exit=0
=== and the CUSTOMER verification copies are untouched ===
bookstack
calibre-web
paperless-ngx
=== the verdict persisted ===
=== the log line ===
2026/08/31 19:10:14 scheduler.go:67: [DEBUG] [scheduler] daily job registered: name="offsite-proof" schedule="05:30" nextRun=2026-09-01T05:30:00+02:00 totalJobs=14
2026/08/31 19:10:14 scheduler.go:67: [DEBUG] [scheduler] daily job offsite-proof: next run at 2026-09-01 05:30:00 CEST (waiting 8h19m45s)
2026/08/31 19:10:42 auth.go:134: [DEBUG] [web] auth: valid session for POST /api/debug/backup/offsite-proof
2026/08/31 19:10:50 offbox_proof.go:175: [INFO] [offbox] proof: bookstack PASSED on snapshot 91154be7 in 2.913s — the backup holds what this app should have
proved_snapshots {'bookstack': '91154be7'}
last_proof_run '2026-08-31T19:10:42Z'
last_proof_stack 'bookstack'
last_proof_snapshot '91154be7'
last_proof_result 'pass'
last_proof_reason '<ABSENT>'
@@ -0,0 +1,10 @@
--- run 2 ---
"duration_ms":2402
"snapshot":"ad2ae49c"
"stack":"calibre-web"
"verdict":"pass"
--- run 3 ---
"duration_ms":3631
"snapshot":"a07c36a1"
"stack":"docmost"
"verdict":"pass"
@@ -0,0 +1,25 @@
=== POSITIVE CONTROL: the sampler must SEE a lock. Run the product integrity check. ===
integrity http=200 wall=42.298944s
=== NOW the proof ===
"stack":"kimai"
"verdict":"pass"
=== SAMPLES ===
19:12:47 locks=0
19:12:51 locks=0
19:12:55 locks=0
19:12:59 locks=1
19:13:03 locks=1
19:13:07 locks=1
19:13:11 locks=1
19:13:15 locks=1
19:13:19 locks=1
19:13:23 locks=1
19:13:27 locks=1
19:13:31 locks=1
19:13:35 locks=0
19:13:39 locks=0
19:13:43 locks=1
19:13:47 locks=0
19:13:51 locks=0 | restic -r <REPO> -o <CMD> restore e7dcf8fd --target /mnt/felhom-drives/hdd_1/backups/offsite-proof/kimai --include /mnt/sys_drive/felhom-data/backups/primary/ki
19:13:55 locks=0
@@ -0,0 +1,14 @@
=== the proof ALONE, nothing else running ===
no_snapshot":false
"stack":"opengist"
"verdict":"pass"
=== SAMPLES (a lock here can only be the proof) ===
19:14:24 locks=0
19:14:28 locks=0
19:14:32 locks=0
19:14:36 locks=0
19:14:40 locks=0
19:14:44 locks=0 | restic -r <REPO> -o <CMD> restore b5aa8f9b --target /mnt/felhom-drives/hdd_1/backups/offsite-proof/opengist --include /mnt/sys_drive/felhom-data/backups/primary
19:14:48 locks=0
19:14:52 locks=0
@@ -0,0 +1,9 @@
=== the EXACT lookup argv the proof uses, run 6x back to back (~15s of continuous holding if it locks) ===
lookups-done
19:15:15 locks=0
19:15:19 locks=0
19:15:23 locks=0
19:15:27 locks=0
19:15:31 locks=0
19:15:35 locks=0
19:15:39 locks=0
@@ -0,0 +1,5 @@
BEFORE:
db_dumps : []
volume_dumps: ['opengist_opengist_data.tar']
tar bytes : 181248
tar sha256 : 3e26592f4abddd75b1b40c026fe91445edff8fb3df2729248550993b29ef6b64
@@ -0,0 +1,7 @@
1. keep a copy of the tar OUTSIDE the unit so the state is restorable
saved 181248 bytes
2. STOP opengist - a stopped app is one of R-403's own named causes: a failed dump leg
stopped
opengist Exited (0) Less than a second ago
3. remove the volume tar - the unit has now LOST its dump, which is the R-403 starting state
removed
@@ -0,0 +1,18 @@
t0=19:17:22Z
offbox/run http=302
=== offsite run log ===
2026/08/31 19:10:50 offbox_proof.go:175: [INFO] [offbox] proof: bookstack PASSED on snapshot 91154be7 in 2.913s — the backup holds what this app should have
2026/08/31 19:11:34 offbox_proof.go:175: [INFO] [offbox] proof: calibre-web PASSED on snapshot ad2ae49c in 2.402s — the backup holds what this app should have
2026/08/31 19:11:45 offbox_proof.go:175: [INFO] [offbox] proof: docmost PASSED on snapshot a07c36a1 in 3.632s — the backup holds what this app should have
2026/08/31 19:13:34 offbox_integrity.go:306: [INFO] [offbox] integrity: check PASSED in 40s (structure, index, and 100% of the pack data re-read)
2026/08/31 19:13:52 offbox_proof.go:175: [INFO] [offbox] proof: kimai PASSED on snapshot e7dcf8fd in 3.987s — the backup holds what this app should have
2026/08/31 19:14:45 offbox_proof.go:175: [INFO] [offbox] proof: opengist PASSED on snapshot b5aa8f9b in 2.194s — the backup holds what this app should have
2026/08/31 19:17:22 auth.go:134: [DEBUG] [web] auth: valid session for POST /backup/offbox/run
2026/08/31 19:17:22 server.go:393: [DEBUG] [web] ServeHTTP: POST /backup/offbox/run from 172.18.0.4:34134
2026/08/31 19:17:22 offbox.go:904: [INFO] [offbox] backup run started (8 app(s) toggled)
=== the opengist unit NOW ===
db_dumps : []
volume_dumps: ['opengist_opengist_data.tar']
tar bytes : 181248
tar sha256 : 3a0547287639c110ada5c340f27678c16e04237392a5339972e191dc2a7dbd45
@@ -0,0 +1 @@
opengist: Up About a minute (healthy)
@@ -0,0 +1,33 @@
constructed unit at /mnt/felhom-drives/hdd_1/r87-construct/backups/primary/opengist
1082 manifest.json
1750 compose/.felhom.yml
287 compose/app.yaml
1260 compose/docker-compose.yml
manifest declares: db_dumps=[] volume_dumps=[]
compose STILL declares its named volume (the expectation source):
volumes:
opengist_data:
ADDITIVE push: one new snapshot, product tags, no forget and no prune.
unable to create lock in backend: repository is already locked exclusively by PID 5751 on demo-hp by root (UID 0, GID 0)
lock was created at 2026-08-31 19:19:39 (24.48081138s ago)
storage ID 20eb88fd
the `unlock` command can be used to remove stale locks
=== opengist snapshots now (newest last) ===
f0c6ce70 2026-08-28 02:16:53 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist
fa862b02 2026-08-29 02:16:49 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist
a5c5c973 2026-08-30 02:17:08 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist
9280850b 2026-08-31 19:19:10 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist
----------------------------------------------------------------------------------------------------------------------
9 snapshots
Added to the repository: 5.921 KiB (3.633 KiB stored)
processed 4 files, 4.276 KiB in 0:01
snapshot f32e1078 saved
=== newest opengist snapshot is now the CONSTRUCTED hollow one ===
----------------------------------------------------------------------------------------------------------------------------------
f32e1078 2026-08-31 19:20:29 demo-hp felhom-offbox,opengist /mnt/felhom-drives/hdd_1/r87-construct/backups/primary/opengist
----------------------------------------------------------------------------------------------------------------------------------
1 snapshots
@@ -0,0 +1,40 @@
=== events pushed BEFORE (baseline) ===
0
=== THE PROOF ===
t0=19:20:42.183Z
{"data":{"duration_ms":2647,"missing":null,"no_snapshot":false,"reason":"","skip_reason":"","skipped":false,"snapshot":"3cab1b88","stack":"bookstack","verdict":"pass"},"message":"A mentés tartalmazza az alkalmazás adatait","ok":true}
http=200 wall=5.383039s
t1=19:20:47.580Z
run: "stack":"calibre-web" "verdict":"pass"
run: "stack":"docmost" "verdict":"pass"
run: "stack":"kimai" "verdict":"pass"
run: "stack":"opengist" "verdict":"fail"
REACHED OPENGIST
=== the controller log line, verbatim ===
2026/08/31 19:21:16 offbox_proof.go:175: [INFO] [offbox] proof: docmost PASSED on snapshot 9493ed08 in 2.963s — the backup holds what this app should have
2026/08/31 19:21:29 offbox_proof.go:175: [INFO] [offbox] proof: kimai PASSED on snapshot 84332185 in 2.988s — the backup holds what this app should have
2026/08/31 19:21:45 offbox_proof.go:181: [ERROR] [offbox] proof: opengist on snapshot f32e1078 is READABLE AND EMPTY (volumes_expected_none_captured: opengist_data) — the store is not damaged; the backup does not contain this app's data
=== the EVENT: how many, what severity ===
2026/08/31 19:21:45 notifier.go:206: [DEBUG] PushEvent: type=offsite_proof_empty severity=error url=https://hub.felhom.eu/api/v1/event
2026/08/31 19:21:45 notifier.go:232: [DEBUG] PushEvent: offsite_proof_empty pushed OK (HTTP 200)
2026/08/31 19:21:45 notifier.go:234: [INFO] Event pushed: offsite_proof_empty (error) — A(z) opengist legutóbbi távoli mentése olvasható, de nem tartalmazza az alkalmazás adatait. A tároló nem sérült — a mentés készült el üresen. A mentést újra el kell készíteni; addig ebből a mentésből nem lehet visszaállítani.
count: 3
=== the persisted verdict ===
python3: can't open file '/tmp/chk.py': [Errno 2] No such file or directory
=== the proof scratch must be gone even on a FAIL ===
(empty above = deleted)
=== persisted verdict ===
proved_snapshots {'bookstack': '3cab1b88', 'calibre-web': 'bc4063b1', 'docmost': '9493ed08', 'kimai': '84332185', 'opengist': 'f32e1078'}
last_proof_run '2026-08-31T19:21:30Z'
last_proof_stack 'opengist'
last_proof_snapshot 'f32e1078'
last_proof_result 'fail'
last_proof_reason 'volumes_expected_none_captured'
=== EXACTLY ONE event? count the PUSH, not the log lines ===
1
@@ -0,0 +1,20 @@
=== 1. the product re-backs-up opengist (healthy unit, tar present) ===
db_dumps : []
volume_dumps: ['opengist_opengist_data.tar']
tar bytes : 181248
tar sha256 : 3a0547287639c110ada5c340f27678c16e04237392a5339972e191dc2a7dbd45
t0=19:22:19Z
offbox/run http=302
run finished after ~15s
=== opengist newest snapshot is HEALTHY again ===
----------------------------------------------------------------------------------------------------------------------------------
f32e1078 2026-08-31 19:20:29 demo-hp felhom-offbox,opengist /mnt/felhom-drives/hdd_1/r87-construct/backups/primary/opengist
----------------------------------------------------------------------------------------------------------------------------------
1 snapshots
opengist newest is no longer the constructed one (after ~80s)
----------------------------------------------------------------------------------------------------------------------
ea94dae0 2026-08-31 19:23:50 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist
----------------------------------------------------------------------------------------------------------------------
1 snapshots
@@ -0,0 +1,7 @@
run finished
"stack":"bookstack" "verdict":"pass"
"stack":"calibre-web" "verdict":"pass"
"stack":"docmost" "verdict":"pass"
"stack":"kimai" "verdict":"pass"
"stack":"opengist" "verdict":"pass"
@@ -0,0 +1,3 @@
=== LIVE Scenario F1 (unplanned, and better than a fixture): the off-site backup run held the
single-writer flag and the proof SKIPPED - duration_ms 0, no verdict, no alarm ===
{"duration_ms":0,"skip_reason":"a backup or restore is already running","skipped":true,"stack":"","verdict":""}
@@ -0,0 +1,40 @@
=== LAYER 3: container /tmp ===
tr: extra operand '"'
Try 'tr --help' for more information.
container /tmp: []
=== LAYER 3b: the constructed tree and the tar backup ===
r87-construct removed
tar backup removed (the live unit has its own healthy tar)
appdata
backups
userdata
=== proof scratch root (empty is correct; the job creates and removes per run) ===
(nothing above = clean)
=== customer verification copies, untouched throughout ===
bookstack
calibre-web
paperless-ngx
=== LAYER 2: guest /root and /tmp ===
tr: extra operand '"'
Try 'tr --help' for more information.
grep: write error: Broken pipe
guest leftovers: []
=== apps healthy ===
opengist Up 3 minutes (healthy)
felhom-controller Up 16 minutes (healthy)
=== teardown re-verified, all three layers ===
--- LAYER 3: container /tmp ---
r87-hollow
(nothing above = empty)
--- LAYER 2: guest ---
.r87-pw.sh
rc=0
--- LAYER 1: PVE host demo-hp ---
rc=1 (1 = clean)
--- container /tmp now ---
--- guest r87 leftovers now ---
grep rc=1 (1 = nothing left)