# SPIKE — can the off-site (restic) copy be restore-tested without a person? (R-87) **Date:** 2026-08-31 · **Venue:** `demo-hp` (Tier 0, disposable), guest 9201, controller **v0.230.0** **Class:** spike. **No production code was written.** No version bump, no build, no deploy, no golden. **Evidence:** `documentation/audits/evidence-spike-restic-restore-2026-08-31/` — 31 files, all pulled off the box before teardown (R-320). **Baselines re-confirmed at the start of the session:** `felhom-controller` `main` @ `2d802d75e88616d86cbade8a0e16965c2b85771c`, `v0.230.0`; `felhom.eu` @ `dddcc808be95d1c89b276b4d791491bad3c96bba` (clean tree, `HEAD == origin/main`); `felhom-agent` `058b945`, `v0.130.0`. --- ## The one-paragraph answer **restic 0.14.0 cannot tell us a restore produced correct files, and neither can anything else the box holds today** — `--verify` exists but checks size and modification time, not content, and the recovery unit's own manifest records a hash for 4 918 bytes of a 213 231 242-byte unit. **But an unattended restore-test is far cheaper than expected:** restoring *every* app on the box — 8 snapshots, 774 MB logical — took **25 seconds**, against **40.3 seconds** for the weekly integrity check sitting beside it. And the question worth asking is not the one R-87 was filed for. Of the five restore-path defects human drills found between 2026-08-26 and 2026-08-31, an unattended scratch-restore would have caught **one**. What it *would* catch, and what the weekly check structurally cannot, is a snapshot that is perfectly intact and contains **nothing recoverable** — the R-403 shape, measured live nine days ago. **Recommendation: build the narrow version (option C below), and do not build the thing R-87 asks for.** --- ## Part 1 — the register correction (done, pushed as `6e550ae`) **When and where it went wrong.** R-87 was moved into `CLOSED-ITEMS.md` by commit **`ef6ac6f`**, 2026-08-22, *"One register, enforced by a gate; closed work compressed into siblings (R-376..R-378)"*. Established from `git log -S`, not inferred: that commit is the only one that ever added an R-87 row to `CLOSED-ITEMS.md` and the only one that removed it from `OPEN-ITEMS.md`. **It is a survivor of R-378, not a separate incident.** R-378 records that same sweep moving six still-open rows — R-123, R-190, R-214, R-264, R-295, R-352 — and restoring them verbatim in the same session. **R-87 is a seventh it missed.** Its state cell read `READY — RE-RANKED UP 2026-08-03 (R-86 closed)`: the leading verdict is `READY`, and the word `closed` later in the same cell describes a *different* row. Nine days in the wrong file, while `OPEN-ITEMS.md`'s ranking paragraph ranked it fourth and pointed at nothing. **The count, reproduced independently — and the predicate decides the answer.** | predicate | rows convicted in `CLOSED-ITEMS.md` (of 151) | which | |---|---|---| | open word anywhere in the **row** | 144 | meaningless | | open word anywhere in the **state cell** | 3 | R-87, **R-224**, **R-260** | | open word in the **leading verdict** | **1** | R-87 | R-224 and R-260 are genuinely closed; their long prose verdicts merely contain the words "open" and "OPEN". **The task author's count of one is right, and it is right only under the leading-verdict predicate** — which is R-378's own lesson, restated by measurement. **The gate:** `scripts/closed_register_gate.py`, two rules — no open state word leading a `CLOSED-ITEMS.md` row's verdict, and no `R-` id with a row in both registers. Red-proofed on both rules; negative-controlled against the files **as they were pushed**, where it convicts R-87 by name and exits 1. Registered as the 12th gate in `repo_gates.py` **after** it was green. Four residual holes are named in its docstring. **One thing the second rule turned up:** `R-398` also had a row in both registers — a deliberate cross-reference stub. It is now prose beneath the table, not a table row. **A duplicate this session did NOT fix:** `OPEN-ITEMS.md` carries two unrelated findings both numbered **R-133** (`:267` hub `customer_id` uniqueness; `:273` plaintext break-glass credential). Filed as **R-406**; deliberately not gated, because a within-register duplicate rule would fail on a pre-existing row and a registered-but-failing gate refuses every push. --- ## Q1 — what restic is actually running? **Answer: restic 0.14.0, and every source comment asserting that is correct.** From the running container on `demo-hp`, not from the Dockerfile: ``` restic 0.14.0 compiled with go1.19.8 on linux/amd64 /usr/bin/restic /etc/debian_version → 12.15 (Debian bookworm, as the Dockerfile says) ``` Method: `pct exec 9201 -- docker exec felhom-controller restic version`, rc=0. Evidence `02-q1-restic-version.txt`. The comments at `offbox_capture.go:15`, `offbox_restore.go:21`, `offbox.go:726`, `offbox_progress.go:48,66` are **confirmed, not corrected**. --- ## Q2 — what can that version verify about a RESTORE? **Answer: `--verify` exists, and it does NOT verify content. It cannot tell us a restore produced correct files.** **`--verify` is real.** `restic restore --help` on the running container lists `--verify verify restored files content`. Controls, because a `grep -c 0` must be earned: | control | result | |---|---| | positive — `--target` present | 1 hit | | `--verify` present | 1 hit, quoted verbatim above | | negative — `--delete`, `--dry-run`, `--overwrite`, `--sparse` (all post-0.14) | 0 hits each | | negative — `ZZZ-NOT-A-FLAG` | 0 hits | Evidence `03-q2-restore-help.txt`, `04-q2-controls.txt`. **`--verify` and `--no-lock` both exist in 0.14.0 and neither appears anywhere in the controller source** — `grep -rn` over `felhom-controller/controller/` returns rc=1 for both, with `--json`/`--target` as the positive control (`05-q2-codebase-verify-grep.txt`). **What `--verify` actually checks — measured, with a red-proof and a negative control.** | test | what was done | result | |---|---|---| | cost | virgin restore of kimai (213 231 242 B, 7 files) with `--verify` | verify itself **131 ms**; restore+verify 3 438 ms | | **red-proof** | corrupt one byte in a restored 160 MB tar, **size and mtime preserved**, re-run `restore --verify` | **PASSED clean, rc=0** — the corruption was not detected | | control (size) | truncate the same file by 1 byte, re-run | restic silently **re-downloaded** it; verify reported "7 files, 133 ms", rc=0 | | control (mtime) | corrupt content and bump mtime, re-run | restic silently **re-downloaded** it; rc=0 | | negative control | same command against the untouched copy | identical output — so "passed" carries no information | **131 ms cannot hash 213 MB.** Combined with the red-proof, `--verify` in 0.14.0 is a size-and-mtime reconciliation that re-fetches anything that disagrees. It is useful — it makes a restore self-repairing — and it is **not** a content check. Evidence `20-`, `21-`, `22-`. **This is the point at which the spike's shape changed**, per §9's instruction to stop and reconsider after Q1/Q2: the tool cannot supply the reference, so Q3 became the hard question. --- ## Q3 — if not restic, then what is the reference for "correct"? **Answer: there is none available to an unattended test today. The one hash record that exists covers 0.002 % of a unit's bytes.** | candidate | verdict | why | |---|---|---| | restic `--verify` | **rejected** | size + mtime only — Q2's red-proof | | the snapshot's own metadata | **rejected** | `restic ls --json` file nodes in 0.14.0 carry name, size, mode, uid/gid and three timestamps — **no content hash** (`24-q3-hash-coverage.txt`) | | a hash the controller already records | **rejected as a content reference, kept as a completeness one** | see below | | `restic check --read-data` | **already shipped, different question** | it proves the STORE's packs, never the restored files | | a planted sentinel | **rejected** | a drill technique. An unattended test may not write data into a customer's app to have something to look for | | the live data | **rejected** | it drifts by design; the snapshot is 12 h old by the time a check runs | | restore twice and compare | **rejected** | proves determinism, not correctness | **The hash record that exists, and its exact coverage.** The recovery unit's `manifest.json` carries a `checksums` object. For kimai, measured on the restored unit: ``` checksums: .felhom.yml (2 235 B), app.yaml (488 B), docker-compose.yml (2 195 B) = 4 918 B files in the unit: + kimai-mariadb.sql 48 217 B + kimai_kimai_db_data.tar 160 331 776 B + kimai_kimai_var.tar 52 845 056 B + manifest.json 1 275 B total 213 231 242 B ``` **4 918 of 213 231 242 bytes — 0.0023 %.** The three config files are hashed; the database dump and the two volume tars, which are the recoverable data, are not. Filed as **R-409**. **What the manifest CAN answer is a different and better question.** It declares `db_dumps` and `volume_dumps` by name, and `unitCarriesData` (`r403_hollow.go:40`) already asks it. A restored unit can therefore be checked for **completeness** — does every file the manifest declares exist — with no new metadata, no new reference, and no content hash. That is the whole of the recommendation in §Recommendation. --- ## Q4 — what does one restore-test cost? **Answer: about 4 seconds per app and 25 seconds for the whole box — cheaper than the weekly integrity check it would sit beside. The cost is dominated by per-snapshot round-trip, not by data volume.** **Through the product's own path** (`POST /backup/offbox/restore` → `RestoreOffboxScratch`), timed from the POST to the completion log line: | app | mode | logical size | wall clock | |---|---|---|---| | docmost | unit | 118 207 270 B | **9 s** (13:40:10 → 13:40:19) | | kimai | full | 213 231 242 B | **11 s** (13:42:54 → 13:43:05) | | opengist | unit | 185 664 B | ~8 s | **Raw restic, all eight snapshots back to back** (`26-q4-all-apps-cost.txt`): | app | logical | restored to disk | ms | |---|---|---|---| | privatebin | 2 123 006 | 2 159 870 | 2 628 | | opengist | 185 664 | 222 528 | 2 253 | | calibre-web | 5 808 703 | 5 878 335 | 3 642 | | paperless-ngx | 83 473 654 | 83 547 382 | 3 062 | | bookstack | 166 641 768 | 166 682 728 | 2 745 | | docmost | 118 207 270 | 118 248 230 | 3 676 | | romm | 184 706 816 | 184 747 776 | 3 978 | | kimai | 213 231 242 | 213 272 202 | 3 202 | | **all eight** | **774 378 123** | **774 759 051** | **25 s** | **185 KB takes 2.25 s and 213 MB takes 3.20 s.** Nearly all of it is fixed per-snapshot cost — opening the repo, loading the index, one SFTP session. Data adds roughly **1 s per 200 MB** (≈ 67 MB/s on this link). **The cache is not hiding the cost.** `/root/.cache/restic` is **1.1 MB** — index and metadata only. Cached 3 198 ms vs `--no-cache` 5 423 ms for kimai; the trees are byte-identical (`diff -r` → YES). Pack data always crosses the wire. Evidence `25-q4-cache-effect.txt`. **Peak scratch:** the restore writes the **full logical size** — 213 272 202 B for the largest app. Sequential-with-cleanup needs only the largest app; all-at-once needs 774 MB. **Against R-359's numbers.** R-359 measured 35.0 s at structure depth and 39.2 s at 100 % read-data on a 140 829 678 B store. Re-measured today through the product's own debug button on a 141 959 062 B store: **40 257 ms at 100 %** (`15-q5-integrity-run.txt`). So: > **restore-testing every app on the box (25 s) costs LESS than one weekly integrity check (40 s).** > Same order of magnitude, and on the cheaper side of it. **Extrapolation — labelled as extrapolation.** R-401 records that the check's cost tracks the index and a read-data run's tracks the data. **A restore tracks the data too, plus a fixed per-snapshot cost.** With 8 apps and the measured ≈ 67 MB/s: | store | fixed (8 × ~2 s) | data | total | |---|---|---|---| | today, 774 MB logical | 16 s | ~9 s | **25 s (measured)** | | 10× — 7.7 GB | 16 s | ~115 s | ~2.2 min | | 100× — 77 GB | 16 s | ~19 min | ~19 min | **Why it may not hold.** One store, one link, one afternoon. This is DooPlex→Hetzner; a customer's domestic line is the real variable, and restore is the *download* direction, usually the faster one on such a line. Scratch space becomes the binding constraint long before time does: at 100× the largest app would need ~21 GB of free scratch, and `RestoreOffboxScratch`'s headroom gate would start refusing. **Unknown, and what would settle it:** a measurement on a store above 10 GB. None exists. This is the same single-data-point hole R-401 already owns. --- ## Q5 — contention **Answer: skip-if-busy stays right, because the operation is seconds and not minutes — but the restore path takes NO single-writer flag at all today, which is a bigger problem than contention.** **The configured window on this box**, read from the scheduler's own registrations rather than assumed (`28-q5-schedule.txt`, times are **CEST** — the guest scheduler runs in local time, not UTC): | job | time | measured duration | |---|---|---| | db-dump | 02:30 | — | | tier2-backup | 03:30 | — | | **offbox-backup** | **04:15** | **2m52s** (last run, `last_duration`) | | offsite-abandon-sweep | 05:10 | — | | **offsite-integrity** | **06:00** | **40.3 s** at 100 % depth | A restore-test of all eight apps holds anything it holds for **25 s** — one seventh of the nightly off-site backup, and shorter than the check beside it. There is an empty gap from ~04:18 to 06:00. **Skip-if-busy remains the right policy**, and it is not the "minutes rather than seconds" case the task worried about. The `integrityCheckTimeout` reasoning transfers unchanged. **The real finding here.** `offbox_integrity.go:28` states the invariant: > *"Every off-site operation takes `acquireRunning` for exactly that reason."* **It does not.** `grep -rn 'acquireRunning()'` finds nine non-test callers; `RestoreOffboxScratch` (`offbox_restore.go:234`) is **not** among them. `restore_wizard.go:174` says the same thing independently — *"`RestoreOffboxScratch` never acquires it at all"* — and the UI works around it with a separate display flag. So a scheduled, unattended caller built on `RestoreOffboxScratch` today would run with **no single-writer flag**, which is precisely the hazard the integrity file's header is shaped around. **A comment asserting an invariant with no test pinning it.** Filed as **R-408**. --- ## Q6 — what does a restore-test WRITE to the repository? **Answer: `restic restore` writes nothing — it does not even take a lock. The product's restore path writes anyway, because it runs `restic unlock` before every restore. And `restic check` — the thing whose comment says it never writes — DOES take a lock.** **Method.** Two observers inside the container, both proven before being believed: 1. a lock sampler running `restic list locks --no-lock` every ~4 s (the observer itself never writes); 2. an argv sampler reading `/proc/*/cmdline` every 0.2 s, with the repo URL and `sftp.command` redacted at the source. **Positive control for the lock sampler — it works.** The product's own integrity check ran 13:41:28 → 13:42:11: ``` 13:41:30 locks=0 13:41:34 locks=1 ids=81fd4d4200d848466e18cb7a9d8e0d4300d95707c68417449178c3023531eecf ... nine consecutive samples, one lock ... 13:42:10 locks=1 ids=81fd… 13:42:14 locks=0 ``` **So `restic check` writes a lock file to the repository.** `offbox_integrity.go:255` says *"It NEVER writes to the repository: `check` is a read verb, and nothing here prunes, forgets, unlocks or backs up."* The three named verbs are correct; the sentence's headline is not. Filed as **R-407**. **The restore, with the same proven instrument:** locks=0 across every sample inside both restore windows (13:40:11/:15/:19 for docmost, 13:42:57/:43:01/:05 for kimai). **restic 0.14.0's `restore` does not lock the repository.** **What the product runs, observed argv (redacted), one restore of `opengist`:** ``` 13:44:41 restic -r sftp: -o sftp.command= snapshots latest --tag opengist --json 13:44:44 restic -r sftp: -o sftp.command= unlock 13:44:46 restic -r sftp: -o sftp.command= restore b5aa8f9b --target … --include … ``` Three invocations. The middle one is `unlockStale` (`offbox_restore.go:289` → `offbox.go:743`), which runs **unconditionally before every restore** and is a **delete verb against `locks/`**. With no stale lock present it deletes nothing — but it is a write-capable command on the exact path R-87 wants to run unattended. `resticStep`'s escalation to `unlock --remove-all` (`offbox.go:760-775`) fires only on `repository is already locked`, which a restore cannot provoke by itself now that we know restore takes no lock. **Against §5's constraint** — *"R-95 still applies: that credential can delete, so a restic restore-test must never be able to write to the repo"*: - **The lead in the task was right in direction and wrong in mechanism.** The write is not `resticStep`'s escalation; it is `unlockStale`, one line earlier and unconditional. - **The constraint IS satisfiable, cheaply, and both mechanisms already exist in 0.14.0 and are unused:** a restore-test that passes `--no-lock` and skips `unlockStale` writes nothing to the repository at all. That is a design note for whoever builds it — **this spike did not fix it**, per §7. --- ## Q7 — what would an unattended restore-test catch that the weekly 100 % check does not? **Answer: of the five defects human drills found in the last six days, one. But that is the wrong scoreboard, and the right one has a much better answer.** | defect | would an unattended scratch-restore have caught it? | how / why not | |---|---|---| | **R-353** — a **local** unit restore reported a bare completion whether it returned a dataset or nothing | **NO** | different code path entirely (`restore_unit.go`). An off-site restore-test never enters it | | **R-354** — the off-site full restore has no named-volume replay leg: the tar reaches the scratch and is never replayed | **NO** | the defect is *after* the scratch. A test that stops at the scratch sees a correct scratch. Going further means replaying into a live app, which an unattended test must not do | | **R-356** — the off-site restore refused every app with no data drive (40 of 53 catalogue templates) | **YES** | the test calls `RestoreOffboxScratch`, gets a refusal, and `rerr != nil`. On this box five of eight apps live on `/mnt/sys_drive` with no drive — it would have fired on the first night | | **R-358** — `OffboxFullScratchReady` asked "non-empty directory", which is what a *failed* restic run leaves | **NO** | needs a restore that fails part-way. On a healthy run the broken and the fixed predicate agree | | **R-403** — a hollow unit mirrored over a complete one with `--delete`; 120 082 104 B → 7 036 B, recorded as success | **NO as filed** — but **YES for the shape** | R-403 destroyed a *local* second-drive copy. What a restore-test sees is the consequence: once a hollow unit is captured off-site, the snapshot is perfectly intact and contains nothing recoverable | **One of five. If the question is "does our restore code work", the honest answer is: the drills already answer it, they answer it better, and automating a worse version of it is not worth an evening.** **The scoreboard that matters is different, and the weekly check structurally cannot play on it.** > `restic check --read-data-subset=100%` proves that **the bytes we stored are the bytes we stored.** > It cannot tell us **we stored the wrong thing.** A hollow recovery unit — no database dump, no volume tar — backs up cleanly, checks cleanly at 100 % depth, restores cleanly, and recovers nothing. **R-403 proved that shape is real, on this fleet, nine days ago, measured in bytes.** Nothing in the product asks the question today, on any tier, at any cadence. A restore-test is simply the cheapest place to ask it, because the manifest that answers it travels inside the snapshot. --- ## Recommendation to Viktor — three options, with costs **Option A — do not build it.** *Cost:* nothing. *What you get:* the weekly 100 % check keeps proving the stored bytes; the restore code keeps being proven by your drills. *What you lose:* nothing that has bitten yet — and the R-403 shape stays invisible until a customer needs the data. **This is a defensible answer** and it is the one R-87's original framing deserves. **Option B — a scheduled attended drill instead.** *Cost:* one of your evenings, monthly. *What you get:* everything a person can see, including the R-354 class that stops at the scratch. *Honest objection:* this is what already happens, and it is what found all five defects. Scheduling it changes nothing except that it now has a date. **Low value for the price.** **Option C — build the NARROW unattended test: one app per night, rotating, restored to scratch, and checked against its own manifest. RECOMMENDED.** *What it does, in one sentence:* restore the newest off-site snapshot of one app into the throwaway scratch, assert that every file the unit's `manifest.json` declares is present, record which snapshot was proved, delete the scratch. *Measured cost, not estimated:* **~4 s and ≤ 213 MB of scratch per night** (one app), or 25 s for all eight. Off-site traffic: one restore's worth, ≈ the app's size. **Less than the weekly integrity check already running beside it.** No new metadata, no new reference, no content hashes — it reuses `unitCarriesData`'s existing manifest read. *What it catches:* R-356 outright, and the R-403 class — a snapshot that is intact and empty — which nothing else in the product asks about. *What it does not catch, stated so nobody expects it to:* R-353, R-354, R-358. Those stay drill work. *Three things it must be built with, all established by this spike:* 1. **`--no-lock`, and skip `unlockStale`** — then it writes nothing to the repository and R-95's constraint is honoured for real (Q6). 2. **Take `acquireRunning`** — `RestoreOffboxScratch` does not, and the whole off-site single-writer story assumes every operation does (Q5, R-408). 3. **Record the SNAPSHOT it proved, not a timestamp** — R-87's own row already says this, and R-86 built the per-archive due-ness model to copy. **If you do nothing:** the weekly check keeps running and keeps being right about the bytes. The first time a hollow unit reaches the off-site store, nothing will notice, and the discovery will be a customer's restore. That is not a hypothetical shape — it is R-403, measured on 2026-08-31. **I would pick C**, scoped exactly as above. It is cheaper than the check beside it, it asks a question nothing else asks, and it needs no invention. **R-87 itself should be RE-SCOPED, not built as written** — from *"restore-test the restic tier"* to *"prove the off-site snapshot still contains a recoverable unit"*. That is your call, so R-87 stays open carrying this verdict. --- ## Register rows opened by this spike | id | what | |---|---| | **R-405** | R-87 was mis-filed by `ef6ac6f`; corrected + gated (CLOSED same session) | | **R-406** | two unrelated findings share the id R-133 in `OPEN-ITEMS.md` | | **R-407** | `restic check` DOES take a repository lock; `offbox_integrity.go:255` says it never writes | | **R-408** | `RestoreOffboxScratch` takes no `acquireRunning`; `offbox_integrity.go:28` asserts every off-site operation does | | **R-409** | the recovery-unit manifest hashes 4 918 B of a 213 231 242 B unit — the data files have no recorded hash | --- ## Probes, teardown and side-effects **All three layers, and two of them really are "nothing was left".** | layer | created | removed | |---|---|---| | PVE host `demo-hp` `/root` | 6 helper scripts + one 0600 password file | all removed; `ls \| grep` returns nothing | | guest 9201 `/root`, `/tmp` | 9 helper/session files + `spike-env.sh` | all removed; grep returns nothing | | container `/tmp` | `spike-env.sh`, two samplers, two logs, two run-flags | all removed; `/tmp` lists empty | **Scratch directories:** four (`docmost`, `kimai`, `privatebin`, `opengist`) were created by the restores this spike drove and were **removed**. Three (`bookstack`, `calibre-web`, `paperless-ngx`) pre-date this session and were **left alone** — this spike never restored them. **Two deliberate state changes on the box, recorded rather than hidden:** 1. The integrity check I ran to control the lock observer **recorded its verdict** — the box's `last_integrity_check` moved to `2026-08-31T13:41:28Z`, `ok=true`, and `last_integrity_depth` moved from `structure` to **`100%`**. Due-ness advanced by seven days. That is the product behaving correctly; it is not a repair and it is not damage. 2. Two logins and four restores appear in the controller log and in the customer-visible restore-op history. **Nothing was written to the off-site repository by hand.** No prune, no forget, no `unlock` issued by me. Every write observed in Q6 was the product's own. **No controller code changed. No golden is owed.** The fleet stays on v0.230.0. The one delivery debt that exists — `golden_currency_gate.py` red because v0.230.0 has no golden — was **already red at `dddcc80`** before this session began, and belongs to the v0.230.0 release, not to this task. --- ## Observations noticed and not acted on - **`paperless-ngx` and `filebrowser` have no off-site snapshot** under the tags the inventory uses — `paperless` returns nothing and `paperless-ngx` returns one; `filebrowser` returns none at all. `filebrowser` is infrastructure, so that may be correct. Not chased. - **`CLOSED-ITEMS.md` has two malformed rows** — R-399 and R-400 supply two columns where the table declares four, so they render with no `Shipped` and no `Evidence`. The new gate prints them as a warning rather than convicting, because an empty state cell is not an open state word. - **Two rows carry a `|` inside their body** (R-309, R-351), which shifts their own cells. Named as the new gate's first residual hole. - **The controller image has no `ps` and no `python3`** — worth knowing before writing any probe that runs inside it. `/proc/*/cmdline` is the substitute that works. - **`git push --no-verify` was used for this session's Part 1 commit** — bypass **#7**, for the reason R-404 exists: a documentation-only push met `golden_currency_gate.py`. Recorded here because R-404 counts them.