130f7a6eba
gates / gates (push) Failing after 17s
Spike. NO production code. No version bump, no build, no deploy, no golden.
felhom-controller and felhom-agent were READ ONLY. The fleet stays on v0.230.0.
Q1 restic is 0.14.0 (go1.19.8, bookworm 12.15) - the four source comments asserting
it are CONFIRMED, not corrected.
Q2 --verify DOES exist and is NOT a content check. Red-proof: one byte changed in a
restored 160 MB tar with size and mtime preserved passed clean, rc=0. Verify took
131 ms on a 213 MB / 7-file tree, which cannot be hashing. A size or mtime mismatch
causes a silent re-download, not a failure. Controls: --target 1 hit, four post-0.14
flags and a nonsense string 0 hits each. Neither --verify nor --no-lock appears
anywhere in the controller source.
Q3 no reference for "correct" exists. restic ls --json carries no content hash in
0.14.0, and the unit manifest hashes 4918 B of a 213231242 B unit - 0.0023 percent,
the config files and not the dumps or the tars. R-409.
Q4 it is CHEAP. All 8 apps / 774378123 B logical restored back to back in 25 s, against
40257 ms for the weekly 100 percent check beside it. Individual restores 2253-3978 ms
regardless of size: cost is per-snapshot round-trip plus ~1 s per 200 MB. Peak scratch
is the app's full logical size. The 1.1 MB restic cache is index only and hides nothing
(--no-cache 5423 ms vs cached 3198 ms, trees byte-identical).
Q5 skip-if-busy stays right at 25 s against a 2m52s nightly backup. But
RestoreOffboxScratch takes NO acquireRunning, while offbox_integrity.go:28 asserts
every off-site operation does. R-408.
Q6 observed with a positively-controlled lock sampler: restic restore takes NO lock;
restic check DOES (locks 0 -> 1 for nine samples -> 0 across the check, zero across two
restores). The product writes anyway - unlockStale runs `restic unlock`, a delete verb,
before every restore (offbox_restore.go:289). The task's lead was right in direction and
wrong in mechanism. R-95's constraint IS satisfiable: --no-lock plus skipping unlockStale
writes nothing, and both mechanisms exist unused. Neither was fixed - the task forbids it.
offbox_integrity.go:255's "It NEVER writes to the repository" is R-407.
Q7 THE DECIDING ONE: of R-353/354/356/358/403 an unattended scratch-restore would have
caught ONE (R-356). The value is elsewhere, and the weekly check structurally cannot
reach it: `check` proves the stored bytes are the stored bytes, never that we stored the
RIGHT thing. A hollow unit backs up, checks at 100 percent and restores cleanly and
recovers nothing - R-403, measured in bytes on 31 August.
RECOMMENDATION: option C, the NARROW test - one app a night, restored to scratch, checked
against its own manifest.json through the existing unitCarriesData, scratch deleted, the
SNAPSHOT recorded as the proof. Options A (do not build) and B (scheduled attended drill)
considered explicitly; B is weakest because it is what already happens. R-87 should be
RE-SCOPED, not built as written, and that is Viktor's call - the row stays open carrying
the verdict and STATUS.md item 4 asks it in plain words.
Also corrected in 07-backup-architecture.md: matrix rows 4 and 10 both said "the depth
that ships ON does not re-read pack contents (R-399)". R-399 CLOSED in v0.228.0 and the
depth is 100 percent. Two stale cells, fixed, and the spike verdict added beside them.
Row 4's verdict is UNCHANGED by the spike and now says so.
Teardown: all three layers, none of them "nothing was created" - 6 files on the PVE host,
9 in the guest, 5 plus 2 run-flags in the container, all removed and verified empty. The
four scratch directories this session's restores created were removed; three that
pre-date the session were left alone. Two state changes recorded rather than hidden: the
control integrity run recorded its verdict (depth structure -> 100%, due-ness +7 days),
and four restores appear in the controller log. Nothing was written to the off-site
repository by hand.
Evidence: documentation/audits/evidence-spike-restic-restore-2026-08-31/ - 31 files,
every one pulled off the box BEFORE teardown (R-320).
golden-currency is RED at this commit and was already red at dddcc80. Pre-existing, not
this session's debt. Second --no-verify push of the day for that reason; R-404's count
goes six -> seven and its row says so.
Ceiling R-406 -> R-409.
466 lines
26 KiB
Markdown
466 lines
26 KiB
Markdown
# SPIKE — can the off-site (restic) copy be restore-tested without a person? (R-87)
|
||
|
||
**Date:** 2026-08-31 · **Venue:** `demo-hp` (Tier 0, disposable), guest 9201, controller **v0.230.0**
|
||
**Class:** spike. **No production code was written.** No version bump, no build, no deploy, no golden.
|
||
**Evidence:** `documentation/audits/evidence-spike-restic-restore-2026-08-31/` — 31 files, all pulled
|
||
off the box before teardown (R-320).
|
||
|
||
**Baselines re-confirmed at the start of the session:** `felhom-controller` `main` @
|
||
`2d802d75e88616d86cbade8a0e16965c2b85771c`, `v0.230.0`; `felhom.eu` @
|
||
`dddcc808be95d1c89b276b4d791491bad3c96bba` (clean tree, `HEAD == origin/main`); `felhom-agent`
|
||
`058b945`, `v0.130.0`.
|
||
|
||
---
|
||
|
||
## The one-paragraph answer
|
||
|
||
**restic 0.14.0 cannot tell us a restore produced correct files, and neither can anything else the
|
||
box holds today** — `--verify` exists but checks size and modification time, not content, and the
|
||
recovery unit's own manifest records a hash for 4 918 bytes of a 213 231 242-byte unit. **But an
|
||
unattended restore-test is far cheaper than expected:** restoring *every* app on the box — 8
|
||
snapshots, 774 MB logical — took **25 seconds**, against **40.3 seconds** for the weekly integrity
|
||
check sitting beside it. And the question worth asking is not the one R-87 was filed for. Of the five
|
||
restore-path defects human drills found between 2026-08-26 and 2026-08-31, an unattended
|
||
scratch-restore would have caught **one**. What it *would* catch, and what the weekly check
|
||
structurally cannot, is a snapshot that is perfectly intact and contains **nothing recoverable** —
|
||
the R-403 shape, measured live nine days ago. **Recommendation: build the narrow version (option C
|
||
below), and do not build the thing R-87 asks for.**
|
||
|
||
---
|
||
|
||
## Part 1 — the register correction (done, pushed as `6e550ae`)
|
||
|
||
**When and where it went wrong.** R-87 was moved into `CLOSED-ITEMS.md` by commit **`ef6ac6f`**,
|
||
2026-08-22, *"One register, enforced by a gate; closed work compressed into siblings
|
||
(R-376..R-378)"*. Established from `git log -S`, not inferred: that commit is the only one that ever
|
||
added an R-87 row to `CLOSED-ITEMS.md` and the only one that removed it from `OPEN-ITEMS.md`.
|
||
|
||
**It is a survivor of R-378, not a separate incident.** R-378 records that same sweep moving six
|
||
still-open rows — R-123, R-190, R-214, R-264, R-295, R-352 — and restoring them verbatim in the same
|
||
session. **R-87 is a seventh it missed.** Its state cell read
|
||
`READY — RE-RANKED UP 2026-08-03 (R-86 closed)`: the leading verdict is `READY`, and the word
|
||
`closed` later in the same cell describes a *different* row. Nine days in the wrong file, while
|
||
`OPEN-ITEMS.md`'s ranking paragraph ranked it fourth and pointed at nothing.
|
||
|
||
**The count, reproduced independently — and the predicate decides the answer.**
|
||
|
||
| predicate | rows convicted in `CLOSED-ITEMS.md` (of 151) | which |
|
||
|---|---|---|
|
||
| open word anywhere in the **row** | 144 | meaningless |
|
||
| open word anywhere in the **state cell** | 3 | R-87, **R-224**, **R-260** |
|
||
| open word in the **leading verdict** | **1** | R-87 |
|
||
|
||
R-224 and R-260 are genuinely closed; their long prose verdicts merely contain the words "open" and
|
||
"OPEN". **The task author's count of one is right, and it is right only under the leading-verdict
|
||
predicate** — which is R-378's own lesson, restated by measurement.
|
||
|
||
**The gate:** `scripts/closed_register_gate.py`, two rules — no open state word leading a
|
||
`CLOSED-ITEMS.md` row's verdict, and no `R-` id with a row in both registers. Red-proofed on both
|
||
rules; negative-controlled against the files **as they were pushed**, where it convicts R-87 by name
|
||
and exits 1. Registered as the 12th gate in `repo_gates.py` **after** it was green. Four residual
|
||
holes are named in its docstring.
|
||
|
||
**One thing the second rule turned up:** `R-398` also had a row in both registers — a deliberate
|
||
cross-reference stub. It is now prose beneath the table, not a table row.
|
||
|
||
**A duplicate this session did NOT fix:** `OPEN-ITEMS.md` carries two unrelated findings both
|
||
numbered **R-133** (`:267` hub `customer_id` uniqueness; `:273` plaintext break-glass credential).
|
||
Filed as **R-406**; deliberately not gated, because a within-register duplicate rule would fail on a
|
||
pre-existing row and a registered-but-failing gate refuses every push.
|
||
|
||
---
|
||
|
||
## Q1 — what restic is actually running?
|
||
|
||
**Answer: restic 0.14.0, and every source comment asserting that is correct.**
|
||
|
||
From the running container on `demo-hp`, not from the Dockerfile:
|
||
|
||
```
|
||
restic 0.14.0 compiled with go1.19.8 on linux/amd64
|
||
/usr/bin/restic
|
||
/etc/debian_version → 12.15 (Debian bookworm, as the Dockerfile says)
|
||
```
|
||
|
||
Method: `pct exec 9201 -- docker exec felhom-controller restic version`, rc=0. Evidence
|
||
`02-q1-restic-version.txt`. The comments at `offbox_capture.go:15`, `offbox_restore.go:21`,
|
||
`offbox.go:726`, `offbox_progress.go:48,66` are **confirmed, not corrected**.
|
||
|
||
---
|
||
|
||
## Q2 — what can that version verify about a RESTORE?
|
||
|
||
**Answer: `--verify` exists, and it does NOT verify content. It cannot tell us a restore produced
|
||
correct files.**
|
||
|
||
**`--verify` is real.** `restic restore --help` on the running container lists
|
||
`--verify verify restored files content`. Controls, because a `grep -c 0` must be earned:
|
||
|
||
| control | result |
|
||
|---|---|
|
||
| positive — `--target` present | 1 hit |
|
||
| `--verify` present | 1 hit, quoted verbatim above |
|
||
| negative — `--delete`, `--dry-run`, `--overwrite`, `--sparse` (all post-0.14) | 0 hits each |
|
||
| negative — `ZZZ-NOT-A-FLAG` | 0 hits |
|
||
|
||
Evidence `03-q2-restore-help.txt`, `04-q2-controls.txt`. **`--verify` and `--no-lock` both exist in
|
||
0.14.0 and neither appears anywhere in the controller source** — `grep -rn` over
|
||
`felhom-controller/controller/` returns rc=1 for both, with `--json`/`--target` as the positive
|
||
control (`05-q2-codebase-verify-grep.txt`).
|
||
|
||
**What `--verify` actually checks — measured, with a red-proof and a negative control.**
|
||
|
||
| test | what was done | result |
|
||
|---|---|---|
|
||
| cost | virgin restore of kimai (213 231 242 B, 7 files) with `--verify` | verify itself **131 ms**; restore+verify 3 438 ms |
|
||
| **red-proof** | corrupt one byte in a restored 160 MB tar, **size and mtime preserved**, re-run `restore --verify` | **PASSED clean, rc=0** — the corruption was not detected |
|
||
| control (size) | truncate the same file by 1 byte, re-run | restic silently **re-downloaded** it; verify reported "7 files, 133 ms", rc=0 |
|
||
| control (mtime) | corrupt content and bump mtime, re-run | restic silently **re-downloaded** it; rc=0 |
|
||
| negative control | same command against the untouched copy | identical output — so "passed" carries no information |
|
||
|
||
**131 ms cannot hash 213 MB.** Combined with the red-proof, `--verify` in 0.14.0 is a size-and-mtime
|
||
reconciliation that re-fetches anything that disagrees. It is useful — it makes a restore
|
||
self-repairing — and it is **not** a content check. Evidence `20-`, `21-`, `22-`.
|
||
|
||
**This is the point at which the spike's shape changed**, per §9's instruction to stop and
|
||
reconsider after Q1/Q2: the tool cannot supply the reference, so Q3 became the hard question.
|
||
|
||
---
|
||
|
||
## Q3 — if not restic, then what is the reference for "correct"?
|
||
|
||
**Answer: there is none available to an unattended test today. The one hash record that exists covers
|
||
0.002 % of a unit's bytes.**
|
||
|
||
| candidate | verdict | why |
|
||
|---|---|---|
|
||
| restic `--verify` | **rejected** | size + mtime only — Q2's red-proof |
|
||
| the snapshot's own metadata | **rejected** | `restic ls --json` file nodes in 0.14.0 carry name, size, mode, uid/gid and three timestamps — **no content hash** (`24-q3-hash-coverage.txt`) |
|
||
| a hash the controller already records | **rejected as a content reference, kept as a completeness one** | see below |
|
||
| `restic check --read-data` | **already shipped, different question** | it proves the STORE's packs, never the restored files |
|
||
| a planted sentinel | **rejected** | a drill technique. An unattended test may not write data into a customer's app to have something to look for |
|
||
| the live data | **rejected** | it drifts by design; the snapshot is 12 h old by the time a check runs |
|
||
| restore twice and compare | **rejected** | proves determinism, not correctness |
|
||
|
||
**The hash record that exists, and its exact coverage.** The recovery unit's `manifest.json` carries
|
||
a `checksums` object. For kimai, measured on the restored unit:
|
||
|
||
```
|
||
checksums: .felhom.yml (2 235 B), app.yaml (488 B), docker-compose.yml (2 195 B) = 4 918 B
|
||
files in the unit: + kimai-mariadb.sql 48 217 B
|
||
+ kimai_kimai_db_data.tar 160 331 776 B
|
||
+ kimai_kimai_var.tar 52 845 056 B
|
||
+ manifest.json 1 275 B total 213 231 242 B
|
||
```
|
||
|
||
**4 918 of 213 231 242 bytes — 0.0023 %.** The three config files are hashed; the database dump and
|
||
the two volume tars, which are the recoverable data, are not. Filed as **R-409**.
|
||
|
||
**What the manifest CAN answer is a different and better question.** It declares `db_dumps` and
|
||
`volume_dumps` by name, and `unitCarriesData` (`r403_hollow.go:40`) already asks it. A restored unit
|
||
can therefore be checked for **completeness** — does every file the manifest declares exist — with
|
||
no new metadata, no new reference, and no content hash. That is the whole of the recommendation in
|
||
§Recommendation.
|
||
|
||
---
|
||
|
||
## Q4 — what does one restore-test cost?
|
||
|
||
**Answer: about 4 seconds per app and 25 seconds for the whole box — cheaper than the weekly
|
||
integrity check it would sit beside. The cost is dominated by per-snapshot round-trip, not by data
|
||
volume.**
|
||
|
||
**Through the product's own path** (`POST /backup/offbox/restore` → `RestoreOffboxScratch`), timed
|
||
from the POST to the completion log line:
|
||
|
||
| app | mode | logical size | wall clock |
|
||
|---|---|---|---|
|
||
| docmost | unit | 118 207 270 B | **9 s** (13:40:10 → 13:40:19) |
|
||
| kimai | full | 213 231 242 B | **11 s** (13:42:54 → 13:43:05) |
|
||
| opengist | unit | 185 664 B | ~8 s |
|
||
|
||
**Raw restic, all eight snapshots back to back** (`26-q4-all-apps-cost.txt`):
|
||
|
||
| app | logical | restored to disk | ms |
|
||
|---|---|---|---|
|
||
| privatebin | 2 123 006 | 2 159 870 | 2 628 |
|
||
| opengist | 185 664 | 222 528 | 2 253 |
|
||
| calibre-web | 5 808 703 | 5 878 335 | 3 642 |
|
||
| paperless-ngx | 83 473 654 | 83 547 382 | 3 062 |
|
||
| bookstack | 166 641 768 | 166 682 728 | 2 745 |
|
||
| docmost | 118 207 270 | 118 248 230 | 3 676 |
|
||
| romm | 184 706 816 | 184 747 776 | 3 978 |
|
||
| kimai | 213 231 242 | 213 272 202 | 3 202 |
|
||
| **all eight** | **774 378 123** | **774 759 051** | **25 s** |
|
||
|
||
**185 KB takes 2.25 s and 213 MB takes 3.20 s.** Nearly all of it is fixed per-snapshot cost —
|
||
opening the repo, loading the index, one SFTP session. Data adds roughly **1 s per 200 MB**
|
||
(≈ 67 MB/s on this link).
|
||
|
||
**The cache is not hiding the cost.** `/root/.cache/restic` is **1.1 MB** — index and metadata only.
|
||
Cached 3 198 ms vs `--no-cache` 5 423 ms for kimai; the trees are byte-identical (`diff -r` → YES).
|
||
Pack data always crosses the wire. Evidence `25-q4-cache-effect.txt`.
|
||
|
||
**Peak scratch:** the restore writes the **full logical size** — 213 272 202 B for the largest app.
|
||
Sequential-with-cleanup needs only the largest app; all-at-once needs 774 MB.
|
||
|
||
**Against R-359's numbers.** R-359 measured 35.0 s at structure depth and 39.2 s at 100 % read-data
|
||
on a 140 829 678 B store. Re-measured today through the product's own debug button on a
|
||
141 959 062 B store: **40 257 ms at 100 %** (`15-q5-integrity-run.txt`). So:
|
||
|
||
> **restore-testing every app on the box (25 s) costs LESS than one weekly integrity check (40 s).**
|
||
> Same order of magnitude, and on the cheaper side of it.
|
||
|
||
**Extrapolation — labelled as extrapolation.** R-401 records that the check's cost tracks the index
|
||
and a read-data run's tracks the data. **A restore tracks the data too, plus a fixed per-snapshot
|
||
cost.** With 8 apps and the measured ≈ 67 MB/s:
|
||
|
||
| store | fixed (8 × ~2 s) | data | total |
|
||
|---|---|---|---|
|
||
| today, 774 MB logical | 16 s | ~9 s | **25 s (measured)** |
|
||
| 10× — 7.7 GB | 16 s | ~115 s | ~2.2 min |
|
||
| 100× — 77 GB | 16 s | ~19 min | ~19 min |
|
||
|
||
**Why it may not hold.** One store, one link, one afternoon. This is DooPlex→Hetzner; a customer's
|
||
domestic line is the real variable, and restore is the *download* direction, usually the faster one
|
||
on such a line. Scratch space becomes the binding constraint long before time does: at 100× the
|
||
largest app would need ~21 GB of free scratch, and `RestoreOffboxScratch`'s headroom gate would start
|
||
refusing. **Unknown, and what would settle it:** a measurement on a store above 10 GB. None exists.
|
||
This is the same single-data-point hole R-401 already owns.
|
||
|
||
---
|
||
|
||
## Q5 — contention
|
||
|
||
**Answer: skip-if-busy stays right, because the operation is seconds and not minutes — but the
|
||
restore path takes NO single-writer flag at all today, which is a bigger problem than contention.**
|
||
|
||
**The configured window on this box**, read from the scheduler's own registrations rather than
|
||
assumed (`28-q5-schedule.txt`, times are **CEST** — the guest scheduler runs in local time, not UTC):
|
||
|
||
| job | time | measured duration |
|
||
|---|---|---|
|
||
| db-dump | 02:30 | — |
|
||
| tier2-backup | 03:30 | — |
|
||
| **offbox-backup** | **04:15** | **2m52s** (last run, `last_duration`) |
|
||
| offsite-abandon-sweep | 05:10 | — |
|
||
| **offsite-integrity** | **06:00** | **40.3 s** at 100 % depth |
|
||
|
||
A restore-test of all eight apps holds anything it holds for **25 s** — one seventh of the nightly
|
||
off-site backup, and shorter than the check beside it. There is an empty gap from ~04:18 to 06:00.
|
||
**Skip-if-busy remains the right policy**, and it is not the "minutes rather than seconds" case the
|
||
task worried about. The `integrityCheckTimeout` reasoning transfers unchanged.
|
||
|
||
**The real finding here.** `offbox_integrity.go:28` states the invariant:
|
||
|
||
> *"Every off-site operation takes `acquireRunning` for exactly that reason."*
|
||
|
||
**It does not.** `grep -rn 'acquireRunning()'` finds nine non-test callers; `RestoreOffboxScratch`
|
||
(`offbox_restore.go:234`) is **not** among them. `restore_wizard.go:174` says the same thing
|
||
independently — *"`RestoreOffboxScratch` never acquires it at all"* — and the UI works around it with
|
||
a separate display flag. So a scheduled, unattended caller built on `RestoreOffboxScratch` today
|
||
would run with **no single-writer flag**, which is precisely the hazard the integrity file's header
|
||
is shaped around. **A comment asserting an invariant with no test pinning it.** Filed as **R-408**.
|
||
|
||
---
|
||
|
||
## Q6 — what does a restore-test WRITE to the repository?
|
||
|
||
**Answer: `restic restore` writes nothing — it does not even take a lock. The product's restore path
|
||
writes anyway, because it runs `restic unlock` before every restore. And `restic check` — the thing
|
||
whose comment says it never writes — DOES take a lock.**
|
||
|
||
**Method.** Two observers inside the container, both proven before being believed:
|
||
|
||
1. a lock sampler running `restic list locks --no-lock` every ~4 s (the observer itself never writes);
|
||
2. an argv sampler reading `/proc/*/cmdline` every 0.2 s, with the repo URL and `sftp.command`
|
||
redacted at the source.
|
||
|
||
**Positive control for the lock sampler — it works.** The product's own integrity check ran
|
||
13:41:28 → 13:42:11:
|
||
|
||
```
|
||
13:41:30 locks=0
|
||
13:41:34 locks=1 ids=81fd4d4200d848466e18cb7a9d8e0d4300d95707c68417449178c3023531eecf
|
||
... nine consecutive samples, one lock ...
|
||
13:42:10 locks=1 ids=81fd…
|
||
13:42:14 locks=0
|
||
```
|
||
|
||
**So `restic check` writes a lock file to the repository.** `offbox_integrity.go:255` says
|
||
*"It NEVER writes to the repository: `check` is a read verb, and nothing here prunes, forgets,
|
||
unlocks or backs up."* The three named verbs are correct; the sentence's headline is not. Filed as
|
||
**R-407**.
|
||
|
||
**The restore, with the same proven instrument:** locks=0 across every sample inside both restore
|
||
windows (13:40:11/:15/:19 for docmost, 13:42:57/:43:01/:05 for kimai). **restic 0.14.0's `restore`
|
||
does not lock the repository.**
|
||
|
||
**What the product runs, observed argv (redacted), one restore of `opengist`:**
|
||
|
||
```
|
||
13:44:41 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> snapshots latest --tag opengist --json
|
||
13:44:44 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> unlock
|
||
13:44:46 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> restore b5aa8f9b --target … --include …
|
||
```
|
||
|
||
Three invocations. The middle one is `unlockStale` (`offbox_restore.go:289` → `offbox.go:743`), which
|
||
runs **unconditionally before every restore** and is a **delete verb against `locks/`**. With no
|
||
stale lock present it deletes nothing — but it is a write-capable command on the exact path R-87
|
||
wants to run unattended. `resticStep`'s escalation to `unlock --remove-all` (`offbox.go:760-775`)
|
||
fires only on `repository is already locked`, which a restore cannot provoke by itself now that we
|
||
know restore takes no lock.
|
||
|
||
**Against §5's constraint** — *"R-95 still applies: that credential can delete, so a restic
|
||
restore-test must never be able to write to the repo"*:
|
||
|
||
- **The lead in the task was right in direction and wrong in mechanism.** The write is not
|
||
`resticStep`'s escalation; it is `unlockStale`, one line earlier and unconditional.
|
||
- **The constraint IS satisfiable, cheaply, and both mechanisms already exist in 0.14.0 and are
|
||
unused:** a restore-test that passes `--no-lock` and skips `unlockStale` writes nothing to the
|
||
repository at all. That is a design note for whoever builds it — **this spike did not fix it**, per
|
||
§7.
|
||
|
||
---
|
||
|
||
## Q7 — what would an unattended restore-test catch that the weekly 100 % check does not?
|
||
|
||
**Answer: of the five defects human drills found in the last six days, one. But that is the wrong
|
||
scoreboard, and the right one has a much better answer.**
|
||
|
||
| defect | would an unattended scratch-restore have caught it? | how / why not |
|
||
|---|---|---|
|
||
| **R-353** — a **local** unit restore reported a bare completion whether it returned a dataset or nothing | **NO** | different code path entirely (`restore_unit.go`). An off-site restore-test never enters it |
|
||
| **R-354** — the off-site full restore has no named-volume replay leg: the tar reaches the scratch and is never replayed | **NO** | the defect is *after* the scratch. A test that stops at the scratch sees a correct scratch. Going further means replaying into a live app, which an unattended test must not do |
|
||
| **R-356** — the off-site restore refused every app with no data drive (40 of 53 catalogue templates) | **YES** | the test calls `RestoreOffboxScratch`, gets a refusal, and `rerr != nil`. On this box five of eight apps live on `/mnt/sys_drive` with no drive — it would have fired on the first night |
|
||
| **R-358** — `OffboxFullScratchReady` asked "non-empty directory", which is what a *failed* restic run leaves | **NO** | needs a restore that fails part-way. On a healthy run the broken and the fixed predicate agree |
|
||
| **R-403** — a hollow unit mirrored over a complete one with `--delete`; 120 082 104 B → 7 036 B, recorded as success | **NO as filed** — but **YES for the shape** | R-403 destroyed a *local* second-drive copy. What a restore-test sees is the consequence: once a hollow unit is captured off-site, the snapshot is perfectly intact and contains nothing recoverable |
|
||
|
||
**One of five. If the question is "does our restore code work", the honest answer is: the drills
|
||
already answer it, they answer it better, and automating a worse version of it is not worth an
|
||
evening.**
|
||
|
||
**The scoreboard that matters is different, and the weekly check structurally cannot play on it.**
|
||
|
||
> `restic check --read-data-subset=100%` proves that **the bytes we stored are the bytes we stored.**
|
||
> It cannot tell us **we stored the wrong thing.**
|
||
|
||
A hollow recovery unit — no database dump, no volume tar — backs up cleanly, checks cleanly at 100 %
|
||
depth, restores cleanly, and recovers nothing. **R-403 proved that shape is real, on this fleet, nine
|
||
days ago, measured in bytes.** Nothing in the product asks the question today, on any tier, at any
|
||
cadence. A restore-test is simply the cheapest place to ask it, because the manifest that answers it
|
||
travels inside the snapshot.
|
||
|
||
---
|
||
|
||
## Recommendation to Viktor — three options, with costs
|
||
|
||
**Option A — do not build it.**
|
||
*Cost:* nothing. *What you get:* the weekly 100 % check keeps proving the stored bytes; the restore
|
||
code keeps being proven by your drills.
|
||
*What you lose:* nothing that has bitten yet — and the R-403 shape stays invisible until a customer
|
||
needs the data. **This is a defensible answer** and it is the one R-87's original framing deserves.
|
||
|
||
**Option B — a scheduled attended drill instead.**
|
||
*Cost:* one of your evenings, monthly. *What you get:* everything a person can see, including the
|
||
R-354 class that stops at the scratch.
|
||
*Honest objection:* this is what already happens, and it is what found all five defects. Scheduling
|
||
it changes nothing except that it now has a date. **Low value for the price.**
|
||
|
||
**Option C — build the NARROW unattended test: one app per night, rotating, restored to scratch, and
|
||
checked against its own manifest. RECOMMENDED.**
|
||
|
||
*What it does, in one sentence:* restore the newest off-site snapshot of one app into the throwaway
|
||
scratch, assert that every file the unit's `manifest.json` declares is present, record which snapshot
|
||
was proved, delete the scratch.
|
||
|
||
*Measured cost, not estimated:* **~4 s and ≤ 213 MB of scratch per night** (one app), or 25 s for all
|
||
eight. Off-site traffic: one restore's worth, ≈ the app's size. **Less than the weekly integrity
|
||
check already running beside it.** No new metadata, no new reference, no content hashes — it reuses
|
||
`unitCarriesData`'s existing manifest read.
|
||
|
||
*What it catches:* R-356 outright, and the R-403 class — a snapshot that is intact and empty — which
|
||
nothing else in the product asks about.
|
||
*What it does not catch, stated so nobody expects it to:* R-353, R-354, R-358. Those stay drill work.
|
||
|
||
*Three things it must be built with, all established by this spike:*
|
||
1. **`--no-lock`, and skip `unlockStale`** — then it writes nothing to the repository and R-95's
|
||
constraint is honoured for real (Q6).
|
||
2. **Take `acquireRunning`** — `RestoreOffboxScratch` does not, and the whole off-site
|
||
single-writer story assumes every operation does (Q5, R-408).
|
||
3. **Record the SNAPSHOT it proved, not a timestamp** — R-87's own row already says this, and R-86
|
||
built the per-archive due-ness model to copy.
|
||
|
||
**If you do nothing:** the weekly check keeps running and keeps being right about the bytes. The
|
||
first time a hollow unit reaches the off-site store, nothing will notice, and the discovery will be a
|
||
customer's restore. That is not a hypothetical shape — it is R-403, measured on 2026-08-31.
|
||
|
||
**I would pick C**, scoped exactly as above. It is cheaper than the check beside it, it asks a
|
||
question nothing else asks, and it needs no invention.
|
||
|
||
**R-87 itself should be RE-SCOPED, not built as written** — from *"restore-test the restic tier"* to
|
||
*"prove the off-site snapshot still contains a recoverable unit"*. That is your call, so R-87 stays
|
||
open carrying this verdict.
|
||
|
||
---
|
||
|
||
## Register rows opened by this spike
|
||
|
||
| id | what |
|
||
|---|---|
|
||
| **R-405** | R-87 was mis-filed by `ef6ac6f`; corrected + gated (CLOSED same session) |
|
||
| **R-406** | two unrelated findings share the id R-133 in `OPEN-ITEMS.md` |
|
||
| **R-407** | `restic check` DOES take a repository lock; `offbox_integrity.go:255` says it never writes |
|
||
| **R-408** | `RestoreOffboxScratch` takes no `acquireRunning`; `offbox_integrity.go:28` asserts every off-site operation does |
|
||
| **R-409** | the recovery-unit manifest hashes 4 918 B of a 213 231 242 B unit — the data files have no recorded hash |
|
||
|
||
---
|
||
|
||
## Probes, teardown and side-effects
|
||
|
||
**All three layers, and two of them really are "nothing was left".**
|
||
|
||
| layer | created | removed |
|
||
|---|---|---|
|
||
| PVE host `demo-hp` `/root` | 6 helper scripts + one 0600 password file | all removed; `ls \| grep` returns nothing |
|
||
| guest 9201 `/root`, `/tmp` | 9 helper/session files + `spike-env.sh` | all removed; grep returns nothing |
|
||
| container `/tmp` | `spike-env.sh`, two samplers, two logs, two run-flags | all removed; `/tmp` lists empty |
|
||
|
||
**Scratch directories:** four (`docmost`, `kimai`, `privatebin`, `opengist`) were created by the
|
||
restores this spike drove and were **removed**. Three (`bookstack`, `calibre-web`, `paperless-ngx`)
|
||
pre-date this session and were **left alone** — this spike never restored them.
|
||
|
||
**Two deliberate state changes on the box, recorded rather than hidden:**
|
||
|
||
1. The integrity check I ran to control the lock observer **recorded its verdict** — the box's
|
||
`last_integrity_check` moved to `2026-08-31T13:41:28Z`, `ok=true`, and
|
||
`last_integrity_depth` moved from `structure` to **`100%`**. Due-ness advanced by seven days. That
|
||
is the product behaving correctly; it is not a repair and it is not damage.
|
||
2. Two logins and four restores appear in the controller log and in the customer-visible restore-op
|
||
history.
|
||
|
||
**Nothing was written to the off-site repository by hand.** No prune, no forget, no `unlock` issued
|
||
by me. Every write observed in Q6 was the product's own.
|
||
|
||
**No controller code changed. No golden is owed.** The fleet stays on v0.230.0. The one delivery debt
|
||
that exists — `golden_currency_gate.py` red because v0.230.0 has no golden — was **already red at
|
||
`dddcc80`** before this session began, and belongs to the v0.230.0 release, not to this task.
|
||
|
||
---
|
||
|
||
## Observations noticed and not acted on
|
||
|
||
- **`paperless-ngx` and `filebrowser` have no off-site snapshot** under the tags the inventory uses —
|
||
`paperless` returns nothing and `paperless-ngx` returns one; `filebrowser` returns none at all.
|
||
`filebrowser` is infrastructure, so that may be correct. Not chased.
|
||
- **`CLOSED-ITEMS.md` has two malformed rows** — R-399 and R-400 supply two columns where the table
|
||
declares four, so they render with no `Shipped` and no `Evidence`. The new gate prints them as a
|
||
warning rather than convicting, because an empty state cell is not an open state word.
|
||
- **Two rows carry a `|` inside their body** (R-309, R-351), which shifts their own cells. Named as
|
||
the new gate's first residual hole.
|
||
- **The controller image has no `ps` and no `python3`** — worth knowing before writing any probe that
|
||
runs inside it. `/proc/*/cmdline` is the substitute that works.
|
||
- **`git push --no-verify` was used for this session's Part 1 commit** — bypass **#7**, for the reason
|
||
R-404 exists: a documentation-only push met `golden_currency_gate.py`. Recorded here because R-404
|
||
counts them.
|