R-87 CLOSED: live evidence, capability row, architecture verdict, registers
gates / gates (push) Failing after 17s

Controller v0.231.0 + hub v0.110.0, both deployed and verified on demo-hp.

LIVE EVIDENCE (documentation/tests/r87-offsite-proof-2026-08-31/, 16 files, endpoint level
through the exact route the debug button invokes):

- THE CASE THAT MATTERS: a hollow unit - compose declaring opengist_data, manifest
  declaring nothing - was pushed to the live store and the proof returned verdict "fail"
  with volumes_expected_none_captured: opengist_data, emitted EXACTLY ONE
  offsite_proof_empty at severity error, and the hub answered HTTP 200. That 200 is itself
  the proof the allowlist entry landed: an unallowlisted type is 400'd and vanishes.
- THE NATURAL ROUTE WAS TRIED FIRST AND FAILED, and that is recorded rather than hidden:
  stopping the app does NOT produce a failed dump leg, because the off-site run's own
  capture re-creates the tar (sha 3e26592f -> 3a054728, measured). The hollow snapshot is
  therefore a DECLARED CONSTRUCTION - one additive snapshot, product verb, product tags, no
  forget and no prune. State restored: the product's own run made a healthy snapshot the
  newest again and the proof then passed opengist.
- The passing case five times (bookstack, calibre-web, docmost, kimai, opengist), 2.2-4.0s
  each, matching the spike's measured band.
- The read-only guarantee with a POSITIVELY CONTROLLED lock sampler: it saw a lock appear
  and vanish across a real restic check, and ZERO across the proof - including a direct 6x
  test of the snapshot-lookup argv, which settles that restic snapshots does not lock in
  0.14.0 either.
- Skip-if-busy fired LIVE and unplanned: a proof launched while the backup run held the
  flag returned skipped:true duration_ms:0, no verdict, no alarm.
- The customer's own verification copies were untouched throughout, which is the safety
  property the separate proof root exists for.

ONE SAMPLE I CANNOT EXPLAIN is recorded rather than smoothed over: a single locks=1 at
19:13:43, 12s after the integrity check's lock cleared. Two independent tests exclude the
proof; I did not establish what it was.

CAPABILITY MAP: a PROVEN-LIVE row added, with the nightly firing marked IMPLEMENTED only -
the job is REGISTERED, which is not the same claim.

07 section 8 MATRIX ROW 4 WAS NOT MOVED, deliberately, and section 10.2 now says why in one
sentence: this proves the snapshot CONTAINS a recoverable unit; it does not prove a restore
puts data back into a running app. Without that sentence the new green tick reads as
covering the drill.

REGISTER: R-87 CLOSED and compressed into CLOSED-ITEMS.md. OPEN 172 -> 171, CLOSED 151 ->
152. No new rows minted. R-408 and R-409 stay open and are referenced by this work.

golden-currency is RED and it is a DECLARED, EXPECTED debt: v0.231.0 is released and the
newest golden carries 0.230.0. The fleet is on 0.230.0; demo-felhom does not have this job.
A golden carrying 0.231.0 is OWED and it is Viktor's call (R-242). This push uses
--no-verify for that reason - bypass #8.
This commit is contained in:
2026-08-31 21:31:56 +02:00
parent 1aeaa30c28
commit 7ee25925f9
22 changed files with 289 additions and 35 deletions
+19 -31
View File
@@ -1,12 +1,12 @@
# STATUS — what works, what's broken, what's next
**Updated 2026-08-31 (third pass) — golden 0.230.0 is baked, vouched and delivered. Both demo
machines are now on the build that stops a good backup copy being deleted; `demo-felhom` moved
itself. Nothing is waiting on you about that release any more.**
**Updated 2026-08-31 (fourth pass) — the box now checks, every night, that one app's remote
backup still has that app's data in it. It caught a deliberately emptied backup on the first try.
One thing is waiting on you: a golden carrying 0.231.0.**
**Earlier 2026-08-31 — I measured whether the box could test its own off-site
restore without you. It can, and it is cheap — but not in the shape we had written down, so
there is a decision for you in item 4. No product code changed.**
**Earlier 2026-08-31 — I measured whether the box could test its own remote restore without you,
then built the narrow version you picked. The measurement is why it is 3 seconds a night and
not an evening's work.**
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
@@ -18,19 +18,21 @@ there is a decision for you in item 4. No product code changed.**
*This section is allowed to be longer than one screen, and each item says what happens if you do
nothing.*
1. **Nothing is waiting on you about the 0.230.0 release.** Golden **0.230.0** is baked,
published and vouched, and the fleet floor is raised to 0.230.0. **`demo-felhom` was still
running 0.229.0 — the build that deletes a good copy — and moved itself across, unattended, in
about two minutes.** Both machines are healthy on 0.230.0, and a machine installed from scratch
now gets the fix too. Reversible if it ever needs to be: re-select the old values and save.
1. **A golden carrying 0.231.0 is owed.** Golden **0.230.0** is baked, vouched and on both
machines. **0.231.0 is the build with the new nightly backup-content check**, and it is on
`demo-hp` only. **If you do nothing:** the check stays on one machine; a newly installed box
does not get it, and neither does `demo-felhom`. Nothing breaks — this adds a check, it does not
fix a defect. Baking and vouching 0.231.0 and raising the floor closes it, the same three-field
change as before.
2. **Nothing else about this release.** Everything in 0.230.0 is a fix to code that ships in the
controller image; no customer action, no data migration, no credential change.
2. **Nothing else about this release.** Everything in 0.231.0 ships in the controller image plus
two register lines in the hub (already live). No customer action, no data migration, no
credential change.
3. **Whether a documents-only push should still be checked for a missing golden** (R-404). We have now
skipped that check **seven times**, each time for a written reason: it runs on every push to the
skipped that check **eight times**, each time for a written reason: it runs on every push to the
website/documentation repository, including pushes that change nothing a machine installs.
**A guard we correctly skip seven times is teaching us to skip it.**
**A guard we correctly skip eight times is teaching us to skip it.**
**The case for narrowing it:** a documents-only push cannot be the one that finishes a release, so
only checking pushes that touch real code would fire on exactly the risky ones and end the habit.
**The case against:** the check was earned — a release went out while machines were still being
@@ -39,25 +41,11 @@ nothing.*
**If you do nothing:** nothing breaks, the skipping stays routine, and the count keeps rising.
I have NOT changed it; this is yours to decide and mine to build.
4. **Whether to have the box check its own off-site RESTORE every night** (R-87). I measured it
today instead of guessing. It is cheap: restoring **every** app on `demo-hp` — 8 backups, 774 MB —
took **25 seconds**, less than the 40 seconds the weekly check beside it already takes. But it
would catch **one** of the five restore faults we found by hand in the last six days, so the
version the old note asked for is not worth building.
**The version that IS worth building is a different question:** the weekly check proves the stored
bytes are the stored bytes. It cannot tell us we stored the **wrong thing** — an empty recovery
package backs up, checks and restores perfectly and gives the customer nothing back. That is not a
theory; it happened on 31 August (R-403). A nightly check of one app against its own packing list
would catch it and needs nothing new built underneath.
**If you do nothing:** the weekly check keeps being right about the bytes, and the first empty
package will be found by a customer trying to restore.
**My pick:** build the narrow version. **Yours to decide**, and I changed no code today.
5. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
4. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
6. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
5. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
as it misled one by an hour.