18 KiB
REPORT — R-87: the box proves its own off-site copy still holds something (2026-08-31)
Controller v0.231.0 · hub v0.110.0 · both deployed and verified live on demo-hp.
1. Baselines, re-checked at the start
| repo | main @ |
matched the task's stated baseline? |
|---|---|---|
felhom-controller |
2d802d75e88616d86cbade8a0e16965c2b85771c (v0.230.0) |
yes |
felhom.eu |
177c75781e11e24fddbaf73c2159f1efd82479e4 |
yes |
felhom-agent |
058b9450648a359856a4102bea4650e33d8884cd (v0.130.0) |
yes, untouched |
All three trees clean, HEAD == origin/main. MinAgent stays 0.129.0.
2. §5's two answers
5.1 The acceptance rule
Two parts, and part 1 alone is the trap.
- everything the manifest declares is present in the restored unit, and
- the manifest declares what the app is supposed to have.
The spike's own summary — "check it against its own packing list" — is part 1, and taken literally it passes a hollow unit, because a hollow unit declares nothing. That is exactly the shape the job exists to catch. Part 2 is the whole value.
JudgeRestoredUnit (r403_hollow.go) returns three outcomes: pass, fail, cannot_judge.
5.2 Where the expectation comes from — and the volume half WAS built
From inside the unit, never from the live box. The snapshot may predate the app's current shape,
and GetDockerVolumes (backup.go) enumerates from live Docker, which answers a different question.
| half | source | built? |
|---|---|---|
| database | DBServiceNames(composePath) on the unit's own captured compose — the same discriminator RestoreFromRecoveryUnit uses, so this cannot disagree with the restore path about what an app is |
yes |
| volumes | ParseComposeNamedVolumes(composePath) on the same file |
yes, as an EXISTENCE check |
The volume half was BUILT, not deferred, and it is deliberately not a name match. §5.2 asked me to
establish whether the unit's compose can answer it before building on it. It can — measured, not
assumed, on all eight real units on demo-hp:
bookstack 2 tars / 2 compose volumes opengist 1 / 1
docmost 3 / 3 privatebin 1 / 1
kimai 2 / 2 calibre-web 1 / 1
romm 3 / 3 paperless-ngx 3 / 3
and <stack>_<volume>.tar held in every case. But "held on eight" is not "derivable": volume tars
are <project>_<volume>.tar and ResolveDockerVolumeNames derives the project from
filepath.Base(filepath.Dir(composePath)), which inside a unit is the literal string compose,
not the stack. So the existence question (does the compose declare named volumes → the manifest must
declare at least one tar) needs zero inference and was built; name-level matching needs the
project-prefix inference and was not. R-355's rule: a claim about the app must never be inferred from
a counter. Half a rule that is true beats a whole rule that is invented.
3. Files created / modified
Created: internal/backup/offbox_proof.go, internal/backup/r87_judgement_test.go,
internal/backup/r87_proof_job_test.go, internal/backup/r87_wiring_test.go,
internal/notify/r87_proof_event_test.go.
Modified (controller): internal/backup/r403_hollow.go (the judgement, beside the R-403
predicate it reuses), internal/backup/offbox_restore.go (offboxScratchDirIn +
unitOnlyHeadroom), internal/backup/offbox_verify_copies.go (offsiteProofRootFor),
internal/backup/offbox.go (report fields), internal/settings/settings.go,
internal/notify/notifier.go, internal/web/handler_debug.go,
internal/web/templates/debug.html, cmd/controller/main.go, plus CHANGELOG.md, CONTEXT.md,
REUSE.md, controller/README.md.
Modified (felhom.eu): hub/internal/api/handler.go, hub/internal/notify/dispatcher.go,
hub/CHANGELOG.md, manifests/hub.yaml, scripts/wire_contract_gate.py, the capability map,
07-backup-architecture.md, ROADMAP.md, both registers.
4. Commits pushed to main
| repo | commit | what |
|---|---|---|
felhom.eu |
1aeaa30 |
hub v0.110.0 — allowlist + operator-only + the wire-contract allowlist |
felhom-controller |
e43b5ec |
v0.231.0 — the judgement, the job, the alarm, 33 tests |
felhom-controller |
303129e |
the by-hand trigger |
felhom.eu |
(this report's commit) | evidence, capability map, architecture, registers |
The hub shipped FIRST and that ordering is load-bearing: an event type the hub does not allowlist is answered 400 and vanishes. Deploying the controller first would have made every alarm silent until the hub caught up.
5. Tests and the red-proofs
33 new tests, all green. Full controller suite 1689 tests / 28 packages / rc=0; hub suite 18
packages / rc=0. All 13 controller gates and all 13 felhom.eu gates OK except the declared
golden-currency debt (§8).
Five red-proofs, run and recorded — each with the wrong value visible:
| # | what was broken | result |
|---|---|---|
| A2 | the rule replaced by "every declared file is present" (§5.1's trap) | TestR87_HollowUnitForAnAppWithADatabaseFAILS FAILED: got verdict="pass" — and two siblings failed with it |
| A3 | the expectation source dropped; alarm on any empty unit | TestR87_AppWithNoDatabaseAndNoVolumesPasses FAILED: got verdict="fail" reason="database_expected_none_captured" |
| B2 | the deferred scratch delete removed | TestR87_ScratchIsDeletedOnEveryPath FAILED on all four rows: "…/backups/offsite-proof/kimai still exists" |
| B3 | the restore routed the customer path's way (unlockStale + resticStep, no --no-lock) |
TestR87_NeverWritesToTheRepository FAILED: the proof issued a WRITE verb "unlock", with the argv printed |
| B5 | ProvedSnapshots recording time.Now() |
TestR87_ProvedSnapshotIsRecordedNotATimestamp FAILED (got "2026-08-31T18:36:47Z") and so did the rotation — "night 2 re-picked bookstack", which is the real consequence |
Three further red-proofs on the wiring, because an AST test that matches nothing passes silently:
dropping the sched.Daily registration, emptying its closure, and un-guarding the alarm each failed
TestR87_JobIsRegisteredInMain / TestR87_RunnerAlarmsOnlyOnFail by name.
Every red-proof was reverted and the file confirmed byte-identical afterwards.
Test count: 1656 → 1689 (+33). (An earlier git stash comparison produced a nonsense "545
before" — stashing the tracked edits left the new untracked files behind, so the tree did not build
and packages silently failed to list. Counted directly instead; the instrument was the problem.)
6. Deployed version
demo-hp guest 9201: gitea.dooplex.hu/admin/felhom-controller:0.231.0 Up (healthy)
hub (k3s): gitea.dooplex.hu/admin/felhom-hub:0.110.0 Synced / Healthy
7. The five live validations — endpoint level, on demo-hp
Method: POST /api/debug/backup/offsite-proof, the exact endpoint the debug button invokes. No
browser on DooPlex; only client-side rendering is unexercised.
7.1 The good case — PASS
{"stack":"bookstack","snapshot":"91154be7","verdict":"pass","duration_ms":2913}
[INFO] proof: bookstack PASSED on snapshot 91154be7 in 2.913s
2.913 s against the spike's measured 2.3–4.0 s band. The scratch was gone afterwards, the
verdict persisted (proved_snapshots {'bookstack': '91154be7'}, last_proof_result 'pass'), and the
customer's own verification copies were untouched — including bookstack's, which sat in
backups/offsite-restore/ throughout.
7.2 The case that matters — the hollow backup was CAUGHT
[ERROR] proof: opengist on snapshot f32e1078 is READABLE AND EMPTY
(volumes_expected_none_captured: opengist_data)
— the store is not damaged; the backup does not contain this app's data
Event pushed: offsite_proof_empty (error) — A(z) opengist legutóbbi távoli mentése olvasható,
de nem tartalmazza az alkalmazás adatait. A tároló nem sérült — a mentés készült el üresen.
A mentést újra el kell készíteni; addig ebből a mentésből nem lehet visszaállítani.
PushEvent: offsite_proof_empty pushed OK (HTTP 200)
Exactly one event (grep -c "Event pushed: offsite_proof_empty" → 1), severity error,
and the hub answered HTTP 200 — which is itself the proof the allowlist entry landed, because an
unallowlisted type is 400'd. Verdict persisted with its reason. Scratch deleted on the failure path
too.
HOW THE SHAPE WAS PRODUCED — the natural route was tried FIRST and it failed. I stopped
opengist, on the reasoning that a stopped app is one of R-403's own named causes (a failed dump
leg), and removed its volume tar. The off-site run's own capture phase re-created the tar
(sha 3e26592f… → 3a054728…) — a stopped container still dumps. So:
DECLARED CONSTRUCTION. The hollow unit was built by hand: the real compose copied verbatim (so it still declares
opengist_data, the expectation source) with a manifest declaringdb_dumps: []andvolume_dumps: []— the manifest a capture writes for a unit that lost its dumps. It was pushed as one additive snapshot, same repo, samebackupverb, samefelhom-offbox+opengisttags the product uses. No forget, no prune, nothing deleted. The healthy history stayed. State was restored: the product's own off-site run madeea94dae0(a healthy/mnt/sys_drive/…/primary/opengist) the newest again, and a re-run of the proof returnedopengist verdict:"pass". The constructed tree was removed.
7.3 The healthy control — PASS, five times
bookstack, calibre-web, docmost, kimai, opengist all passed, 2.2–4.0 s each.
calibre-web, opengist and privatebin have no database at all, so the no-database branch of
Scenario C is proven live, not only in a fixture.
What is NOT live-proven, stated rather than glossed: no app on
demo-hphas neither a database nor a named volume, so the exact neither/nor instance of Scenario C has no live subject. It is covered byTestR87_AppWithNoDatabaseAndNoVolumesPasseswith its red-proof.
7.4 The read-only proof — no lock, no write verb
The lock sampler was positively controlled before it was believed: across a real
restic check it went locks=0 → locks=1 for nine consecutive samples → locks=0. Against the
proof, run in isolation:
19:14:36 locks=0
19:14:40 locks=0
19:14:44 locks=0 | restic … restore b5aa8f9b --target …/offsite-proof/opengist --include …
19:14:48 locks=0
Zero locks, with the restore caught in flight. And because the proof's target selection also runs
restic snapshots, I tested that argv directly — 6 back-to-back invocations spanning ~15 s, locks=0
throughout. restic snapshots does not lock in 0.14.0 either.
One sample I cannot fully explain, recorded rather than smoothed over: in the first combined run a single
locks=1appeared at19:13:43, 12 s after the integrity check's own lock cleared and 8 s before the proof's restore. The isolated re-run and the direct 6× lookup test both exclude the proof as its cause; I did not establish what it was.
7.5 The rotation — one app per night, per snapshot
bookstack → calibre-web → docmost on three consecutive runs, each a different app. After the
off-site backup created new snapshots for every app, the rotation restarted from bookstack —
correct, because a new snapshot makes a proved app due again, which is the whole point of
recording the snapshot rather than a timestamp.
7.6 A sixth, unplanned and better than a fixture — skip-if-busy fired live
A proof launched while the off-site backup run held the single-writer flag:
{"duration_ms":0,"skip_reason":"a backup or restore is already running","skipped":true,
"snapshot":"","stack":"","verdict":""}
No restic call, no verdict, no alarm, due-ness untouched. Scenario F1 on real hardware.
8. The schedule slot, and the live times it was chosen from
05:30, read off the running box rather than a document:
| job | slot | measured |
|---|---|---|
| db-dump | 02:30 | — |
| tier2-backup | 03:30 | — |
| offbox-backup | 04:15 | 2m52s |
| offsite-abandon-sweep | 05:10 | — |
| offsite-proof | 05:30 ← new | one app 2.2–4.0 s |
| offsite-integrity | 06:00 | 40.3 s |
Confirmed registered on the box:
Daily job offsite-proof scheduled for 2026-09-01 05:30:00 CEST (waiting 8h19m45s), totalJobs=14.
9. NOT yet live-validated — explicit
- The unattended nightly firing. The job is REGISTERED; that is not the same claim. First real firing 2026-09-01 05:30 CEST.
- The fleet.
demo-felhomis on 0.230.0 and does not have this job. Onlydemo-hpwas deployed. - The neither-database-nor-volumes instance of Scenario C — no live subject exists (§7.3).
- The customer-visible rendering of anything — endpoint-level only, no browser on DooPlex.
- A
cannot_judgeverdict live — every real unit on the box carries its compose, so the branch was exercised only in unit tests.
10. Capability map, and what did NOT move
Added: a PROVEN-LIVE row for "the box proves its own off-site copy still holds something",
citing documentation/tests/r87-offsite-proof-2026-08-31/, with the scheduled firing marked
IMPLEMENTED only.
07-backup-architecture.md §8 matrix row 4 was NOT moved, deliberately. This proves the snapshot
contains a recoverable unit; it does not prove a restore puts data back into a running app. §10.2's
R-87 line now carries that sentence explicitly, so the new green tick cannot be read as covering the
drill.
11. Teardown — all three layers
| layer | created | after |
|---|---|---|
PVE host demo-hp /root |
5 scripts + one 0600 password file | ls | grep → nothing |
guest 9201 /root, /tmp |
11 files | grep → nothing (two leftovers found on the first pass and removed) |
container /tmp |
env, sampler, log, run-flag, constructed tree | ls -A /tmp → empty |
The off-site store holds one extra opengist snapshot (f32e1078, the declared construction).
Nothing was deleted from it. Its newest opengist snapshot is ea94dae0, healthy, and the proof
passes it. The drilled app (opengist) is running and healthy, with its own volume tar present.
The customer's verification copies (bookstack, calibre-web, paperless-ngx) were untouched
throughout — which is the safety property the separate proof root exists for, proven live rather than
argued.
12. Register
| id | action |
|---|---|
| R-87 | CLOSED — shipped + proven-live, then compressed into CLOSED-ITEMS.md |
| R-242 | unchanged — a golden carrying 0.231.0 is now owed |
| R-408, R-409 | unchanged and still open; both are referenced by this work and neither was fixed |
Register size, counted from git rather than from memory: OPEN-ITEMS.md 172 → 171 rows (R-87 moved out); CLOSED-ITEMS.md 151 → 152. R-87 now appears exactly once, in CLOSED-ITEMS.md, and closed_register_gate.py confirms it is not in both.
No new rows were minted — every gap this session found is either fixed here or already has a row.
golden-currency is RED and that is a DECLARED, EXPECTED debt: controller v0.231.0 is released
and the newest golden carries 0.230.0. The fleet is on 0.230.0. A golden carrying 0.231.0 is
OWED, and it is Viktor's call (R-242). Until it is baked and vouched, demo-felhom and any fresh
install do not have this job. Both pushes of the felhom.eu repo used --no-verify for that reason
and it is declared here.
13. Scope: two things beyond the task's list, both deliberate
felhom.eu/hub/was touched. The task listeddocumentation/andSTATUS.mdonly. Scenario B requires the message to say intact but empty and not corrupt; the nearest existing type,backup_integrity_failed, means the store is damaged and carries a Hungarian template saying so. Reusing it would have shipped the wrong sentence. Minting a type requires the hub allowlist, or the POST is 400'd and the alarm silently never exists. Reasoned inCONTEXT.mdruling 4.- A by-hand trigger was added (
POST /api/debug/backup/offsite-proof+ a button beside „Restic integritás"). Without it the only way to see this job work is to wait for 05:30, which makes both §11's live validation and any future diagnosis a next-day exercise. Same function as the scheduled job — no second code path.
Observations
- A STOPPED app still dumps its volume. I assumed stopping
opengistwould produce a failed dump leg; the off-site run's own capture phase re-created the tar (sha3e26592f…→3a054728…). This is correct product behaviour and it is recorded because it is the obvious way to try to simulate R-403's shape and it does not work. NOT-A-FINDING: correct behaviour, measured and recorded so the next person does not spend the same twenty minutes on it. - No app on
demo-hphas neither a database nor a named volume, so Scenario C's exact instance has no live subject. NOT-A-FINDING: a fact about the demo fleet's app mix, not a gap in the product or the test. - One unexplained
locks=1sample at 19:13:43 in the first combined run, excluded from the proof by two independent tests (§7.4). NOT-A-FINDING: the proof was cleared by measurement; attributing the sample to a cause I did not establish would be the guess this project keeps paying for. ResolveDockerVolumeNamescannot be used from inside a recovery unit — it derives the compose project from the file's parent directory, which inside a unit is the literal stringcompose. This is why the volume half is an existence check. NOT-A-FINDING: a documented consequence of the unit's layout, recorded inREUSE.mdandCONTEXT.mdwhere the next reader will meet it.