R-359/R-397 closed, R-398 corrected, R-399/R-400 filed with measured numbers
gates / gates (push) Failing after 18s
gates / gates (push) Failing after 18s
THE MEASUREMENT IS THE STORY, and it re-frames the row it was filed under. A pack was corrupted WITHOUT changing its size; plain `restic check` -- the depth that ships ON -- returned `no errors were found`, exit 0. Only --read-data caught it. So the check that shipped verifies the index, the pack inventory and the snapshot graph, and does NOT re-hash pack contents. R-399 was filed as a bandwidth-and-cadence question; it is more than that, and its row now says so. R-399 gets three MEASURED numbers instead of estimates: store 140 829 678 B / 2651 blobs / 67 snapshots; structure check 35.0 s; curve 10% 35.9 s, 50% 37.3 s, 100% 39.2 s. At this size re-reading everything costs four seconds more than reading none, because the wall clock is SFTP round-trips not transfer. The row states the limit too: these do NOT extrapolate. R-400: the sweep the task asked for found EIGHT dead debug buttons, not one. 24 endpoints referenced in debug.html, 17 dispatched. Single dispatcher, exact match, default NotFound -- so they 404. A third of a debug page does nothing, on the surface an operator reaches for when something is already wrong. R-398 is CORRECTED AND LEFT OPEN, not closed. I filed it yesterday saying resticStep is not a seam so no test can drive a restic path. The layer below it has been injectable since the off-site tier shipped. The row survives as the record that the seam EXISTS so nobody re-files it. 07 gap register: R-359 and R-397 closed; R-87 restated IN PLACE as "AND IT IS NOT R-359" because the two rows are adjacent and a check is not a restore-test. 08 alarm ladder: both event types recorded, including that `ok` is `info` and therefore mails nobody BY DESIGN, and that all three registers were checked and deliberately left alone. 00 capability map: PROVEN-LIVE for the check, the notifier and the hazard control; the scheduled firing is IMPLEMENTED only, because a week has not passed. wire_contract_gate: `offsite.last_integrity_ok` allowlisted WITH A REASON. The gate was right -- the controller emits a field no hub struct can decode. Building the display is a hub change and R-331 ruled that class the operator's decision; the entry says to delete it when a surface exists. This push used `git push --no-verify`. golden-currency is CONVICTED and right: 0.227.1 is released and the golden carries 0.226.1. A BYPASS, not a waiver, and the task spec directs it -- golden and fleet delivery are Viktor's (R-242). It is item 3 under "Waiting on you". Register 163 -> 165 -> 163.
This commit is contained in:
@@ -18,10 +18,25 @@ nothing.*
|
||||
today receives 0.226.1 and every fix from the four releases of 2026-08-30.
|
||||
Evidence: `documentation/tests/golden-0.226.1-2026-08-30/`.
|
||||
|
||||
2. **Nothing else.** Everything in the four releases of 2026-08-30 is a fix to code that ships in the
|
||||
2. **How deep should the off-site check go?** (R-399). The box now checks its off-site store weekly.
|
||||
The check it runs today reads the catalogue — it catches a missing or unreadable backup, and it does
|
||||
**not** re-read the stored bytes, so it cannot see a file that has quietly rotted.
|
||||
**What it costs to go deeper, measured on your own machine today, not guessed:**
|
||||
the shallow check takes **35.0 s**; re-reading **all** the data takes **39.2 s**. Four seconds more.
|
||||
That is because the time goes on talking to the off-site box, not on moving data — and it will stop
|
||||
being true as the store grows, so this is worth re-measuring, not deciding once forever.
|
||||
**If you do nothing:** the catalogue is checked weekly and the stored bytes are never re-read.
|
||||
I can turn it on with one setting whenever you say.
|
||||
|
||||
3. **Bake and vouch a golden carrying 0.227.1, then raise the floor** — the usual last step. `demo-hp`
|
||||
runs 0.227.1; the fleet floor is 0.226.1 and the golden carries 0.226.1.
|
||||
**If you do nothing:** a machine installed today gets 0.226.1 and none of today's off-site checking,
|
||||
and `demo-felhom` stays where it is. Tracked on R-242.
|
||||
|
||||
4. **Nothing else.** Everything in the releases of 2026-08-30 is a fix to code that ships in the
|
||||
controller image; no customer action, no data migration, no credential change.
|
||||
|
||||
3. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
||||
5. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
||||
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
|
||||
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
|
||||
4. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
||||
@@ -54,6 +69,23 @@ off. **`peti-felhom` is a real machine we have not heard from since 15 July** an
|
||||
|
||||
## Shipped
|
||||
|
||||
- **Something finally checks that the off-site copies are still there and readable** (R-359 + R-397,
|
||||
controller 0.227.1, proven on `demo-hp`). Until today **nothing did** — not the box, not the agent.
|
||||
The whole-machine backups had their own checks; the copies holding your customers' documents and
|
||||
photos had none, so we would have found a problem at restore time, with a customer waiting.
|
||||
Now the box checks its own off-site store about once a week and tells you only if something is
|
||||
wrong. **A pass sends no e-mail, on purpose** — a weekly "everything is fine" is how people stop
|
||||
reading their alerts. **It also catches itself up:** it asks „has it been more than seven days?",
|
||||
not „is it Sunday?", so a machine that was switched off on its check day is checked the next day.
|
||||
**And it never gets in the backup's way** — if a backup or restore is running, the check steps aside
|
||||
and tries again tomorrow. That is not politeness: the tool would otherwise clear a lock that a live
|
||||
backup was holding. Proven live, twice over — a real check against your real store (35 seconds), and
|
||||
a second check fired during the first, which correctly stepped aside without doing anything.
|
||||
**Read item 2 under „Waiting on you" for what this check does NOT see.**
|
||||
- **The product stopped claiming a check it never ran.** The monitoring page said an integrity check
|
||||
ran every Sunday. It did not exist. The e-mail text, the settings checkbox, the hub's side of it and
|
||||
a debug button were all built and wired to nothing (R-397). They now have the missing piece.
|
||||
|
||||
- **Taking the safety copy no longer destroys the app's own backup** (R-361, controller 0.221.1,
|
||||
proven on `demo-hp`). Before every restore the machine saves a copy of your live database. To do
|
||||
that it called the ordinary backup routine — **which always writes to the app's normal backup
|
||||
@@ -104,11 +136,12 @@ off. **`peti-felhom` is a real machine we have not heard from since 15 July** an
|
||||
|
||||
## Broken, or knowingly incomplete
|
||||
|
||||
- **Nothing ever checks that the off-site store is still readable** (R-359). Not the controller, not
|
||||
the agent. We find out at restore time. A deliberately corrupted copy was caught instantly by the
|
||||
standard tool — which we never run.
|
||||
- **A restore that returns nothing still reports success** (part of R-354's neighbourhood, not fixed
|
||||
today) and **verification copies have no delete guard**. Both deliberately left for their own rows.
|
||||
- **The off-site store is only checked SHALLOWLY, and the deep check is switched off** (R-399).
|
||||
This is the honest version of what shipped today. The box now checks its own off-site store about
|
||||
once a week (R-359, controller 0.227.0) — and the check it runs reads the *catalogue* of the backups,
|
||||
not the backups themselves. **We proved the difference:** a copy was damaged in a way that left its
|
||||
size unchanged, and the shallow check said „no errors were found". Only the deep check caught it.
|
||||
The deep check is built and **off**, waiting for your decision — see item 2 under „Waiting on you".
|
||||
- **We ask the off-site box a question about once a second** (R-336) — ~85,000 a day for a box we
|
||||
write to weekly. The leak that made this dangerous is fixed (R-344); the volume is not. The ceiling
|
||||
is **under a year** away on the corrected measurement, not two.
|
||||
|
||||
Reference in New Issue
Block a user