R-87 SPIKE: measured, do not build it as written (R-407..R-409 filed)
gates / gates (push) Failing after 17s
gates / gates (push) Failing after 17s
Spike. NO production code. No version bump, no build, no deploy, no golden.
felhom-controller and felhom-agent were READ ONLY. The fleet stays on v0.230.0.
Q1 restic is 0.14.0 (go1.19.8, bookworm 12.15) - the four source comments asserting
it are CONFIRMED, not corrected.
Q2 --verify DOES exist and is NOT a content check. Red-proof: one byte changed in a
restored 160 MB tar with size and mtime preserved passed clean, rc=0. Verify took
131 ms on a 213 MB / 7-file tree, which cannot be hashing. A size or mtime mismatch
causes a silent re-download, not a failure. Controls: --target 1 hit, four post-0.14
flags and a nonsense string 0 hits each. Neither --verify nor --no-lock appears
anywhere in the controller source.
Q3 no reference for "correct" exists. restic ls --json carries no content hash in
0.14.0, and the unit manifest hashes 4918 B of a 213231242 B unit - 0.0023 percent,
the config files and not the dumps or the tars. R-409.
Q4 it is CHEAP. All 8 apps / 774378123 B logical restored back to back in 25 s, against
40257 ms for the weekly 100 percent check beside it. Individual restores 2253-3978 ms
regardless of size: cost is per-snapshot round-trip plus ~1 s per 200 MB. Peak scratch
is the app's full logical size. The 1.1 MB restic cache is index only and hides nothing
(--no-cache 5423 ms vs cached 3198 ms, trees byte-identical).
Q5 skip-if-busy stays right at 25 s against a 2m52s nightly backup. But
RestoreOffboxScratch takes NO acquireRunning, while offbox_integrity.go:28 asserts
every off-site operation does. R-408.
Q6 observed with a positively-controlled lock sampler: restic restore takes NO lock;
restic check DOES (locks 0 -> 1 for nine samples -> 0 across the check, zero across two
restores). The product writes anyway - unlockStale runs `restic unlock`, a delete verb,
before every restore (offbox_restore.go:289). The task's lead was right in direction and
wrong in mechanism. R-95's constraint IS satisfiable: --no-lock plus skipping unlockStale
writes nothing, and both mechanisms exist unused. Neither was fixed - the task forbids it.
offbox_integrity.go:255's "It NEVER writes to the repository" is R-407.
Q7 THE DECIDING ONE: of R-353/354/356/358/403 an unattended scratch-restore would have
caught ONE (R-356). The value is elsewhere, and the weekly check structurally cannot
reach it: `check` proves the stored bytes are the stored bytes, never that we stored the
RIGHT thing. A hollow unit backs up, checks at 100 percent and restores cleanly and
recovers nothing - R-403, measured in bytes on 31 August.
RECOMMENDATION: option C, the NARROW test - one app a night, restored to scratch, checked
against its own manifest.json through the existing unitCarriesData, scratch deleted, the
SNAPSHOT recorded as the proof. Options A (do not build) and B (scheduled attended drill)
considered explicitly; B is weakest because it is what already happens. R-87 should be
RE-SCOPED, not built as written, and that is Viktor's call - the row stays open carrying
the verdict and STATUS.md item 4 asks it in plain words.
Also corrected in 07-backup-architecture.md: matrix rows 4 and 10 both said "the depth
that ships ON does not re-read pack contents (R-399)". R-399 CLOSED in v0.228.0 and the
depth is 100 percent. Two stale cells, fixed, and the spike verdict added beside them.
Row 4's verdict is UNCHANGED by the spike and now says so.
Teardown: all three layers, none of them "nothing was created" - 6 files on the PVE host,
9 in the guest, 5 plus 2 run-flags in the container, all removed and verified empty. The
four scratch directories this session's restores created were removed; three that
pre-date the session were left alone. Two state changes recorded rather than hidden: the
control integrity run recorded its verdict (depth structure -> 100%, due-ness +7 days),
and four restores appear in the controller log. Nothing was written to the off-site
repository by hand.
Evidence: documentation/audits/evidence-spike-restic-restore-2026-08-31/ - 31 files,
every one pulled off the box BEFORE teardown (R-320).
golden-currency is RED at this commit and was already red at dddcc80. Pre-existing, not
this session's debt. Second --no-verify push of the day for that reason; R-404's count
goes six -> seven and its row says so.
Ceiling R-406 -> R-409.
This commit is contained in:
@@ -1,6 +1,10 @@
|
||||
# STATUS — what works, what's broken, what's next
|
||||
|
||||
**Updated 2026-08-31 — the weekly off-site check now re-reads your actual data, not just the list of
|
||||
**Updated 2026-08-31 (second pass) — I measured whether the box could test its own off-site
|
||||
restore without you. It can, and it is cheap — but not in the shape we had written down, so
|
||||
there is a decision for you in item 4. No code changed today.**
|
||||
|
||||
**Earlier 2026-08-31 — the weekly off-site check now re-reads your actual data, not just the list of
|
||||
it. A third of the debug page did nothing and no longer exists. 0.228.0 is baked, vouched and
|
||||
delivered; both demo machines are on it and nothing is waiting on you about this release.**
|
||||
|
||||
@@ -24,9 +28,9 @@ nothing.*
|
||||
controller image; no customer action, no data migration, no credential change.
|
||||
|
||||
3. **Whether a documents-only push should still be checked for a missing golden** (R-404). We have now
|
||||
skipped that check **six times**, each time for a written reason: it runs on every push to the
|
||||
skipped that check **seven times**, each time for a written reason: it runs on every push to the
|
||||
website/documentation repository, including pushes that change nothing a machine installs.
|
||||
**A guard we correctly skip six times is teaching us to skip it.**
|
||||
**A guard we correctly skip seven times is teaching us to skip it.**
|
||||
**The case for narrowing it:** a documents-only push cannot be the one that finishes a release, so
|
||||
only checking pushes that touch real code would fire on exactly the risky ones and end the habit.
|
||||
**The case against:** the check was earned — a release went out while machines were still being
|
||||
@@ -35,11 +39,25 @@ nothing.*
|
||||
**If you do nothing:** nothing breaks, the skipping stays routine, and the count keeps rising.
|
||||
I have NOT changed it; this is yours to decide and mine to build.
|
||||
|
||||
4. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
||||
4. **Whether to have the box check its own off-site RESTORE every night** (R-87). I measured it
|
||||
today instead of guessing. It is cheap: restoring **every** app on `demo-hp` — 8 backups, 774 MB —
|
||||
took **25 seconds**, less than the 40 seconds the weekly check beside it already takes. But it
|
||||
would catch **one** of the five restore faults we found by hand in the last six days, so the
|
||||
version the old note asked for is not worth building.
|
||||
**The version that IS worth building is a different question:** the weekly check proves the stored
|
||||
bytes are the stored bytes. It cannot tell us we stored the **wrong thing** — an empty recovery
|
||||
package backs up, checks and restores perfectly and gives the customer nothing back. That is not a
|
||||
theory; it happened on 31 August (R-403). A nightly check of one app against its own packing list
|
||||
would catch it and needs nothing new built underneath.
|
||||
**If you do nothing:** the weekly check keeps being right about the bytes, and the first empty
|
||||
package will be found by a customer trying to restore.
|
||||
**My pick:** build the narrow version. **Yours to decide**, and I changed no code today.
|
||||
|
||||
5. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
||||
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
|
||||
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
|
||||
|
||||
5. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
||||
6. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
||||
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
|
||||
as it misled one by an hour.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user