R-87 SPIKE: measured, do not build it as written (R-407..R-409 filed)
gates / gates (push) Failing after 17s
gates / gates (push) Failing after 17s
Spike. NO production code. No version bump, no build, no deploy, no golden.
felhom-controller and felhom-agent were READ ONLY. The fleet stays on v0.230.0.
Q1 restic is 0.14.0 (go1.19.8, bookworm 12.15) - the four source comments asserting
it are CONFIRMED, not corrected.
Q2 --verify DOES exist and is NOT a content check. Red-proof: one byte changed in a
restored 160 MB tar with size and mtime preserved passed clean, rc=0. Verify took
131 ms on a 213 MB / 7-file tree, which cannot be hashing. A size or mtime mismatch
causes a silent re-download, not a failure. Controls: --target 1 hit, four post-0.14
flags and a nonsense string 0 hits each. Neither --verify nor --no-lock appears
anywhere in the controller source.
Q3 no reference for "correct" exists. restic ls --json carries no content hash in
0.14.0, and the unit manifest hashes 4918 B of a 213231242 B unit - 0.0023 percent,
the config files and not the dumps or the tars. R-409.
Q4 it is CHEAP. All 8 apps / 774378123 B logical restored back to back in 25 s, against
40257 ms for the weekly 100 percent check beside it. Individual restores 2253-3978 ms
regardless of size: cost is per-snapshot round-trip plus ~1 s per 200 MB. Peak scratch
is the app's full logical size. The 1.1 MB restic cache is index only and hides nothing
(--no-cache 5423 ms vs cached 3198 ms, trees byte-identical).
Q5 skip-if-busy stays right at 25 s against a 2m52s nightly backup. But
RestoreOffboxScratch takes NO acquireRunning, while offbox_integrity.go:28 asserts
every off-site operation does. R-408.
Q6 observed with a positively-controlled lock sampler: restic restore takes NO lock;
restic check DOES (locks 0 -> 1 for nine samples -> 0 across the check, zero across two
restores). The product writes anyway - unlockStale runs `restic unlock`, a delete verb,
before every restore (offbox_restore.go:289). The task's lead was right in direction and
wrong in mechanism. R-95's constraint IS satisfiable: --no-lock plus skipping unlockStale
writes nothing, and both mechanisms exist unused. Neither was fixed - the task forbids it.
offbox_integrity.go:255's "It NEVER writes to the repository" is R-407.
Q7 THE DECIDING ONE: of R-353/354/356/358/403 an unattended scratch-restore would have
caught ONE (R-356). The value is elsewhere, and the weekly check structurally cannot
reach it: `check` proves the stored bytes are the stored bytes, never that we stored the
RIGHT thing. A hollow unit backs up, checks at 100 percent and restores cleanly and
recovers nothing - R-403, measured in bytes on 31 August.
RECOMMENDATION: option C, the NARROW test - one app a night, restored to scratch, checked
against its own manifest.json through the existing unitCarriesData, scratch deleted, the
SNAPSHOT recorded as the proof. Options A (do not build) and B (scheduled attended drill)
considered explicitly; B is weakest because it is what already happens. R-87 should be
RE-SCOPED, not built as written, and that is Viktor's call - the row stays open carrying
the verdict and STATUS.md item 4 asks it in plain words.
Also corrected in 07-backup-architecture.md: matrix rows 4 and 10 both said "the depth
that ships ON does not re-read pack contents (R-399)". R-399 CLOSED in v0.228.0 and the
depth is 100 percent. Two stale cells, fixed, and the spike verdict added beside them.
Row 4's verdict is UNCHANGED by the spike and now says so.
Teardown: all three layers, none of them "nothing was created" - 6 files on the PVE host,
9 in the guest, 5 plus 2 run-flags in the container, all removed and verified empty. The
four scratch directories this session's restores created were removed; three that
pre-date the session were left alone. Two state changes recorded rather than hidden: the
control integrity run recorded its verdict (depth structure -> 100%, due-ness +7 days),
and four restores appear in the controller log. Nothing was written to the off-site
repository by hand.
Evidence: documentation/audits/evidence-spike-restic-restore-2026-08-31/ - 31 files,
every one pulled off the box BEFORE teardown (R-320).
golden-currency is RED at this commit and was already red at dddcc80. Pre-existing, not
this session's debt. Second --no-verify push of the day for that reason; R-404's count
goes six -> seven and its row says so.
Ceiling R-406 -> R-409.
This commit is contained in:
@@ -0,0 +1,176 @@
|
|||||||
|
# REPORT — R-87 spike: can the off-site copy be restore-tested without a person? (2026-08-31)
|
||||||
|
|
||||||
|
Written as `REPORT-r87-spike.md`, not `REPORT.md`: this repo's `CLAUDE.md` says the shared report is
|
||||||
|
overwritten and two sessions clobber each other.
|
||||||
|
|
||||||
|
**Findings document (the deliverable):** `documentation/audits/SPIKE-restic-restore-test-2026-08-31.md`
|
||||||
|
**Evidence:** `documentation/audits/evidence-spike-restic-restore-2026-08-31/` — 31 files.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. Baselines, re-checked at the start
|
||||||
|
|
||||||
|
| repo | `main` @ | version | matched the task's stated baseline? |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `felhom-controller` | `2d802d75e88616d86cbade8a0e16965c2b85771c` | v0.230.0 | yes |
|
||||||
|
| `felhom.eu` | `dddcc808be95d1c89b276b4d791491bad3c96bba` | — | yes; clean tree, `HEAD == origin/main` |
|
||||||
|
| `felhom-agent` | `058b945` | v0.130.0 | yes |
|
||||||
|
|
||||||
|
## 2. Part 1 — the register correction
|
||||||
|
|
||||||
|
- **The mis-filing commit:** `ef6ac6f`, 2026-08-22, *"One register, enforced by a gate; closed work
|
||||||
|
compressed into siblings (R-376..R-378)"*. Established by `git log -S '**R-87**'` on both register
|
||||||
|
files — it is the only commit that added the row to `CLOSED-ITEMS.md` and the only one that removed
|
||||||
|
it from `OPEN-ITEMS.md`. **The history settled it; no guess was needed.**
|
||||||
|
- **It is a survivor of R-378, not a separate incident.** R-378 records six rows moved wrongly by that
|
||||||
|
same sweep and restored in the same session. R-87 is a **seventh it missed**, and it escaped because
|
||||||
|
its state cell led with `READY` and carried the word `closed` later, describing a different row.
|
||||||
|
- **My own count, reproduced:** **1** mis-filed row under the leading-verdict predicate. The same scan
|
||||||
|
convicts **3** if it reads the whole state cell (R-224 and R-260 are false positives — long prose
|
||||||
|
verdicts containing "open"/"OPEN") and **144** if it reads the whole row. The task author's count of
|
||||||
|
one is confirmed, and only under the predicate R-378 argues for.
|
||||||
|
- **The gate:** `scripts/closed_register_gate.py`. Two rules. **Red-proof rule 1:** a planted `READY`
|
||||||
|
row convicts by name, rc=1; removing it leaves the file byte-identical. **Red-proof rule 2:** a
|
||||||
|
planted duplicate id convicts, rc=1. **Negative control:** run against the files *as pushed*
|
||||||
|
(`HEAD:`), it convicts R-87 at L72 and R-398 at L139, rc=1. **Registered LAST**, as the 12th gate in
|
||||||
|
`repo_gates.py`, after it was green.
|
||||||
|
- **Also corrected:** `R-398` had a row in both registers (a deliberate cross-reference stub). Now
|
||||||
|
prose beneath the table.
|
||||||
|
|
||||||
|
## 3. Q1–Q7, one paragraph each
|
||||||
|
|
||||||
|
**Q1 — restic 0.14.0**, `go1.19.8`, Debian bookworm 12.15, from the running container. The four source
|
||||||
|
comments asserting 0.14.0 are **confirmed**. Method: `docker exec felhom-controller restic version`.
|
||||||
|
|
||||||
|
**Q2 — `--verify` exists and is NOT a content check.** `restic restore --help` lists
|
||||||
|
`--verify verify restored files content`; positive control `--target` = 1 hit, negative controls
|
||||||
|
`--delete/--dry-run/--overwrite/--sparse` and a nonsense string = 0 hits each. Neither `--verify` nor
|
||||||
|
`--no-lock` appears anywhere in the controller source (`grep -rn` rc=1, with `--json`/`--target` as
|
||||||
|
the positive control). **Red-proof:** one byte changed in a restored 160 MB tar with size and mtime
|
||||||
|
preserved — `restore --verify` **passed clean, rc=0**. Verify took **131 ms** on a 213 MB / 7-file
|
||||||
|
tree, which cannot be hashing. A size or mtime mismatch causes a silent **re-download**, not a
|
||||||
|
failure. **So restic cannot tell us a restore produced correct files.**
|
||||||
|
|
||||||
|
**Q3 — no reference for "correct" exists today.** `restic ls --json` file nodes in 0.14.0 carry no
|
||||||
|
content hash. The recovery unit's `manifest.json` hashes three config files — **4 918 B of a
|
||||||
|
213 231 242 B unit, 0.0023 %** — and not the DB dump or the volume tars. A planted sentinel is a drill
|
||||||
|
technique and does not transfer; the live data drifts. **What the manifest CAN answer is
|
||||||
|
completeness**, through the existing `unitCarriesData` (`r403_hollow.go:40`), with no new metadata.
|
||||||
|
Filed as R-409.
|
||||||
|
|
||||||
|
**Q4 — ~4 s per app, 25 s for all eight, cheaper than the weekly check.** Through the product's own
|
||||||
|
path: docmost unit 9 s, kimai full 11 s. Raw restic, all 8 snapshots / **774 378 123 B logical** back
|
||||||
|
to back: **25 s**, individual times 2 253–3 978 ms *regardless of size* (185 KB → 2.25 s, 213 MB →
|
||||||
|
3.20 s). The cost is per-snapshot round-trip plus ≈ 1 s per 200 MB. Peak scratch = the app's full
|
||||||
|
logical size, 213 272 202 B for the largest. The restic cache is **1.1 MB** (index only) and does not
|
||||||
|
hide the cost: `--no-cache` 5 423 ms vs cached 3 198 ms, trees byte-identical. **Against R-359's
|
||||||
|
35.0 s / 39.2 s: the same 100 % check re-measured today is 40 257 ms — so restore-testing the whole
|
||||||
|
box costs LESS than one weekly check.** Extrapolation to 10×/100× is in the findings doc, **labelled
|
||||||
|
as extrapolation**, with scratch space named as the constraint that binds before time does; the
|
||||||
|
single-store hole is R-401's.
|
||||||
|
|
||||||
|
**Q5 — skip-if-busy stays right, and a bigger thing is wrong.** Scheduler registrations read off the
|
||||||
|
box (CEST): db-dump 02:30, tier2 03:30, **offbox-backup 04:15 (2m52s measured)**, abandon-sweep 05:10,
|
||||||
|
**offsite-integrity 06:00 (40.3 s)**. A 25 s hold is seconds, not minutes, and there is an empty gap
|
||||||
|
04:18–06:00. **But `RestoreOffboxScratch` takes no `acquireRunning` at all** — nine non-test callers,
|
||||||
|
it is not one — while `offbox_integrity.go:28` asserts *"Every off-site operation takes
|
||||||
|
`acquireRunning`"*. `restore_wizard.go:174` records the same fact independently. Filed as **R-408**.
|
||||||
|
|
||||||
|
**Q6 — the restore itself writes nothing; the product writes anyway; and `check` writes a lock.**
|
||||||
|
Observed with a lock sampler and an argv sampler, both inside the container, the repo URL redacted at
|
||||||
|
source. **Positive control:** across the product's integrity run the repo went `locks=0` →
|
||||||
|
`locks=1 id=81fd4d42…` for nine consecutive samples → `locks=0`. **The same instrument saw zero locks
|
||||||
|
across two restores**, so `restic restore` in 0.14.0 does not lock. Observed argv for one restore:
|
||||||
|
`snapshots latest --tag <app> --json`, then **`unlock`**, then `restore <id> --target …`. The middle
|
||||||
|
one is `unlockStale` (`offbox_restore.go:289`), unconditional, a **delete verb**. **§5's lead was
|
||||||
|
right in direction and wrong in mechanism** — the write is `unlockStale`, not `resticStep`'s
|
||||||
|
escalation. **The constraint IS satisfiable:** `--no-lock` + skipping `unlockStale` writes nothing,
|
||||||
|
and both mechanisms exist in 0.14.0 unused. `offbox_integrity.go:255`'s *"It NEVER writes to the
|
||||||
|
repository"* is filed as **R-407**. **Neither was fixed** — §7 forbids it.
|
||||||
|
|
||||||
|
**Q7 — one of five.** R-353 (a *local* restore path) **no**; R-354 (no volume-replay leg, after the
|
||||||
|
scratch) **no**; **R-356 (refused every driveless app) YES — five of eight apps on this box would
|
||||||
|
have fired it on the first night**; R-358 (needs a part-way failure) **no**; R-403 (destroyed a
|
||||||
|
*local* copy) **no as filed, yes for the shape**. **The honest verdict is "few", and it points
|
||||||
|
elsewhere:** `check` proves the stored bytes are the stored bytes, never that we stored the *right*
|
||||||
|
thing. A hollow unit backs up, checks at 100 % and restores cleanly, and recovers nothing — R-403,
|
||||||
|
measured in bytes nine days ago. Nothing asks that question on any tier.
|
||||||
|
|
||||||
|
## 4. Recommendation
|
||||||
|
|
||||||
|
Three options with costs and a do-nothing outcome are in the findings document. **I would pick option
|
||||||
|
C — the narrow test:** one app a night, restored to scratch, checked against its own `manifest.json`,
|
||||||
|
scratch deleted, the snapshot recorded as the proof. ~4 s and ≤ 213 MB per night; catches R-356 and
|
||||||
|
the R-403 class; needs no new metadata. **It must use `--no-lock`, skip `unlockStale`, and take
|
||||||
|
`acquireRunning`** — all three established by this spike. **Options A (do not build) and B (scheduled
|
||||||
|
attended drill) were considered explicitly and are argued in the document; B is the weakest, because
|
||||||
|
it is what already happens.** **R-87 should be RE-SCOPED, not built as written — and that is Viktor's
|
||||||
|
call**, so the row stays open carrying the verdict, and `STATUS.md` item 4 asks it in plain words.
|
||||||
|
|
||||||
|
## 5. Evidence
|
||||||
|
|
||||||
|
`documentation/audits/evidence-spike-restic-restore-2026-08-31/`, 31 files, numbered by question.
|
||||||
|
**Every file was pulled off the box before any teardown** (R-320) — including the two in-container
|
||||||
|
sampler logs, which were `cat`ed to DooPlex before the container `/tmp` was cleared.
|
||||||
|
|
||||||
|
## 6. Probes removed — all three layers, and none of them is "nothing was created"
|
||||||
|
|
||||||
|
| layer | created | after teardown |
|
||||||
|
|---|---|---|
|
||||||
|
| PVE host `/root` | 6 scripts + one 0600 password file | `ls \| grep` → nothing |
|
||||||
|
| guest 9201 `/root`, `/tmp` | 9 files | grep → nothing |
|
||||||
|
| container `/tmp` | 5 files + 2 run-flags | `/tmp` lists empty; no restic process left |
|
||||||
|
|
||||||
|
**Scratch directories:** the four created by this session's restores (`docmost`, `kimai`,
|
||||||
|
`privatebin`, `opengist`) were removed. Three (`bookstack`, `calibre-web`, `paperless-ngx`) pre-date
|
||||||
|
this session and were **left alone**. The local password copy was `shred -u`'d.
|
||||||
|
|
||||||
|
**Two state changes recorded rather than hidden:** the integrity check run as the lock positive
|
||||||
|
control **recorded its verdict** (`last_integrity_check` → `2026-08-31T13:41:28Z`, depth `structure` →
|
||||||
|
**`100%`**, due-ness advanced 7 days), and four restores plus two logins appear in the controller log.
|
||||||
|
**Nothing was written to the off-site repository by hand.**
|
||||||
|
|
||||||
|
## 7. Register
|
||||||
|
|
||||||
|
| id | action |
|
||||||
|
|---|---|
|
||||||
|
| **R-87** | **moved back to `OPEN-ITEMS.md`** (verbatim from `ef6ac6f^`, beside R-95), then updated with the spike verdict and a re-scope proposal |
|
||||||
|
| **R-398** | de-tabled in `CLOSED-ITEMS.md`; the open row is the record |
|
||||||
|
| **R-404** | bypass count corrected six → **seven** (this session's Part 1 push) |
|
||||||
|
| **R-405** | filed + **CLOSED** — the mis-file, the reproduced count, the gate |
|
||||||
|
| **R-406** | filed — two findings share the id R-133 |
|
||||||
|
| **R-407** | filed — `restic check` takes a lock; the comment says it never writes |
|
||||||
|
| **R-408** | filed — `RestoreOffboxScratch` takes no `acquireRunning` |
|
||||||
|
| **R-409** | filed — the unit manifest hashes 0.002 % of the unit |
|
||||||
|
|
||||||
|
**Register size:** `OPEN-ITEMS.md` **165 → 171** rows (+R-87 restored, +R-405..R-409); `CLOSED-ITEMS.md` **153 → 151** (−R-87, −R-398).
|
||||||
|
Ceiling **R-404 → R-409**.
|
||||||
|
|
||||||
|
`python3 scripts/unproven.py --summary`: 55 claims, walked 20 / partial 17 / built 14 / missing 4,
|
||||||
|
**NOT WALKED 35 of 55 — unchanged by this session**, which shipped no product claim.
|
||||||
|
|
||||||
|
## 8. No controller code changed and no golden is owed
|
||||||
|
|
||||||
|
`felhom-controller` and `felhom-agent` were **read only**. No version bump, no build, no deploy, no
|
||||||
|
golden. The fleet stays on v0.230.0. **The one golden debt that exists — v0.230.0 released with the
|
||||||
|
newest bake at 0.229.0 — was already red at `dddcc80` before this session started** and belongs to
|
||||||
|
that release, not to this task; `golden_currency_gate.py` was the only failing gate at every point in
|
||||||
|
this session, before and after.
|
||||||
|
|
||||||
|
**`git push --no-verify` was used, twice, for exactly that reason** — records-only pushes meeting the
|
||||||
|
golden gate. That is R-404's subject and the count is updated in its row.
|
||||||
|
|
||||||
|
## 9. Observations — noticed, not acted on
|
||||||
|
|
||||||
|
- `paperless-ngx` has an off-site snapshot under that tag and none under `paperless`; `filebrowser`
|
||||||
|
has none at all (it is infrastructure, so that may be correct). Not chased.
|
||||||
|
- `CLOSED-ITEMS.md` rows **R-399** and **R-400** supply two columns where the table declares four —
|
||||||
|
they render with no `Shipped` and no `Evidence`. The new gate warns rather than convicts, because an
|
||||||
|
empty state cell is not an open state word.
|
||||||
|
- Two rows (**R-309**, **R-351**) carry a `|` inside their body, shifting their own cells. Named as
|
||||||
|
the gate's first residual hole.
|
||||||
|
- The controller image has **no `ps` and no `python3`**. `/proc/*/cmdline` is the substitute that
|
||||||
|
works, and it is worth knowing before writing any probe that runs in there.
|
||||||
|
- The guest scheduler logs in **CEST**, not UTC — `offbox-backup scheduled for 2026-09-01 04:15 CEST`
|
||||||
|
against a `last_run` of `02:17:57Z`. Consistent, and the opposite of what the project memory says
|
||||||
|
about guest time.
|
||||||
@@ -1,6 +1,10 @@
|
|||||||
# STATUS — what works, what's broken, what's next
|
# STATUS — what works, what's broken, what's next
|
||||||
|
|
||||||
**Updated 2026-08-31 — the weekly off-site check now re-reads your actual data, not just the list of
|
**Updated 2026-08-31 (second pass) — I measured whether the box could test its own off-site
|
||||||
|
restore without you. It can, and it is cheap — but not in the shape we had written down, so
|
||||||
|
there is a decision for you in item 4. No code changed today.**
|
||||||
|
|
||||||
|
**Earlier 2026-08-31 — the weekly off-site check now re-reads your actual data, not just the list of
|
||||||
it. A third of the debug page did nothing and no longer exists. 0.228.0 is baked, vouched and
|
it. A third of the debug page did nothing and no longer exists. 0.228.0 is baked, vouched and
|
||||||
delivered; both demo machines are on it and nothing is waiting on you about this release.**
|
delivered; both demo machines are on it and nothing is waiting on you about this release.**
|
||||||
|
|
||||||
@@ -24,9 +28,9 @@ nothing.*
|
|||||||
controller image; no customer action, no data migration, no credential change.
|
controller image; no customer action, no data migration, no credential change.
|
||||||
|
|
||||||
3. **Whether a documents-only push should still be checked for a missing golden** (R-404). We have now
|
3. **Whether a documents-only push should still be checked for a missing golden** (R-404). We have now
|
||||||
skipped that check **six times**, each time for a written reason: it runs on every push to the
|
skipped that check **seven times**, each time for a written reason: it runs on every push to the
|
||||||
website/documentation repository, including pushes that change nothing a machine installs.
|
website/documentation repository, including pushes that change nothing a machine installs.
|
||||||
**A guard we correctly skip six times is teaching us to skip it.**
|
**A guard we correctly skip seven times is teaching us to skip it.**
|
||||||
**The case for narrowing it:** a documents-only push cannot be the one that finishes a release, so
|
**The case for narrowing it:** a documents-only push cannot be the one that finishes a release, so
|
||||||
only checking pushes that touch real code would fire on exactly the risky ones and end the habit.
|
only checking pushes that touch real code would fire on exactly the risky ones and end the habit.
|
||||||
**The case against:** the check was earned — a release went out while machines were still being
|
**The case against:** the check was earned — a release went out while machines were still being
|
||||||
@@ -35,11 +39,25 @@ nothing.*
|
|||||||
**If you do nothing:** nothing breaks, the skipping stays routine, and the count keeps rising.
|
**If you do nothing:** nothing breaks, the skipping stays routine, and the count keeps rising.
|
||||||
I have NOT changed it; this is yours to decide and mine to build.
|
I have NOT changed it; this is yours to decide and mine to build.
|
||||||
|
|
||||||
4. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
4. **Whether to have the box check its own off-site RESTORE every night** (R-87). I measured it
|
||||||
|
today instead of guessing. It is cheap: restoring **every** app on `demo-hp` — 8 backups, 774 MB —
|
||||||
|
took **25 seconds**, less than the 40 seconds the weekly check beside it already takes. But it
|
||||||
|
would catch **one** of the five restore faults we found by hand in the last six days, so the
|
||||||
|
version the old note asked for is not worth building.
|
||||||
|
**The version that IS worth building is a different question:** the weekly check proves the stored
|
||||||
|
bytes are the stored bytes. It cannot tell us we stored the **wrong thing** — an empty recovery
|
||||||
|
package backs up, checks and restores perfectly and gives the customer nothing back. That is not a
|
||||||
|
theory; it happened on 31 August (R-403). A nightly check of one app against its own packing list
|
||||||
|
would catch it and needs nothing new built underneath.
|
||||||
|
**If you do nothing:** the weekly check keeps being right about the bytes, and the first empty
|
||||||
|
package will be found by a customer trying to restore.
|
||||||
|
**My pick:** build the narrow version. **Yours to decide**, and I changed no code today.
|
||||||
|
|
||||||
|
5. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
||||||
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
|
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
|
||||||
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
|
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
|
||||||
|
|
||||||
5. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
6. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
||||||
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
|
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
|
||||||
as it misled one by an hour.
|
as it misled one by an hour.
|
||||||
|
|
||||||
|
|||||||
@@ -98,7 +98,7 @@ with nothing but their dashboard password. No operator, no ticket, no scheduling
|
|||||||
| 1 | „Visszaállítás indítása" | `POST /backup/restore` | destructive — rebuilds the app from its Tier-1 unit |
|
| 1 | „Visszaállítás indítása" | `POST /backup/restore` | destructive — rebuilds the app from its Tier-1 unit |
|
||||||
| 2 | „Fájlok visszaállítása" | `POST /backup/tier2/restore` | additive, missing-only |
|
| 2 | „Fájlok visszaállítása" | `POST /backup/tier2/restore` | additive, missing-only |
|
||||||
| 3 | „Visszaállítás a távoli tárolóból" | `POST /backup/offbox/restore` | non-destructive, to a verification copy |
|
| 3 | „Visszaállítás a távoli tárolóból" | `POST /backup/offbox/restore` | non-destructive, to a verification copy |
|
||||||
| 4 | „Helyreállítás az élő adatok közé" | `POST /backup/offbox/place` | additive, missing-only **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open), and the depth that ships ON does not re-read pack contents (R-399).** |
|
| 4 | „Helyreállítás az élő adatok közé" | `POST /backup/offbox/place` | additive, missing-only **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. |
|
||||||
| 5 | „Teljes visszaállítás (fájlok + adatbázis)" | `POST /backup/offbox/reconstitute` | destructive to files + DB, never deleting |
|
| 5 | „Teljes visszaállítás (fájlok + adatbázis)" | `POST /backup/offbox/reconstitute` | destructive to files + DB, never deleting |
|
||||||
| 6 | „Megosztások visszaállítása" + place | `POST /backup/shares/{restore,place}` | additive |
|
| 6 | „Megosztások visszaállítása" + place | `POST /backup/shares/{restore,place}` | additive |
|
||||||
| 7 | `.fab` import | `POST` → `apiImportStart` | destructive re-import of one app |
|
| 7 | `.fab` import | `POST` → `apiImportStart` | destructive re-import of one app |
|
||||||
@@ -869,7 +869,7 @@ crosses the line — **R-158**.
|
|||||||
| 7 | **Guest stopped and does not come back** | everything | guest-power watchdog (60 s, `onboot` as the deliberate-stop discriminator) | automatic | **120 s** | | **PROVEN** | agent v0.107.0 replay — 120 s unattended vs the incident's 587 s with a human |
|
| 7 | **Guest stopped and does not come back** | everything | guest-power watchdog (60 s, `onboot` as the deliberate-stop discriminator) | automatic | **120 s** | | **PROVEN** | agent v0.107.0 replay — 120 s unattended vs the incident's 587 s with a human |
|
||||||
| 8 | **Host dies (hardware), drives intact** | the data drives; ep0's PBS namespace; the hub's Recipe + Escrow | **install a new host, then** `--selftest=bring-up -mode dr` per guest, **then** re-attach drives by `durable_id` | **operator** (SSH) | | 7 d | **IMPLEMENTED** | bring-up code exists and has **never been executed** (`CAMPAIGN-8…:522`); the drive half of the plan is empty on every box (**R-105**) |
|
| 8 | **Host dies (hardware), drives intact** | the data drives; ep0's PBS namespace; the hub's Recipe + Escrow | **install a new host, then** `--selftest=bring-up -mode dr` per guest, **then** re-attach drives by `durable_id` | **operator** (SSH) | | 7 d | **IMPLEMENTED** | bring-up code exists and has **never been executed** (`CAMPAIGN-8…:522`); the drive half of the plan is empty on every box (**R-105**) |
|
||||||
| 9 | **Whole box lost (fire/theft) — host and drives gone** | ep0 PBS namespace; the restic repo; the hub's Recipe + Escrow | new hardware → day-0 → escrow-consume with **R** → restore guests from PBS → app data from Tier-3 | **operator + customer** (R) | | 7 d (guest) · 24 h (app data) | **IMPLEMENTED / UNPROVEN** | every leg exists; the composed path has never been run. The destructive S5 drill is operator-gated and unrun (`06-offsite-connectivity.md:327`) |
|
| 9 | **Whole box lost (fire/theft) — host and drives gone** | ep0 PBS namespace; the restic repo; the hub's Recipe + Escrow | new hardware → day-0 → escrow-consume with **R** → restore guests from PBS → app data from Tier-3 | **operator + customer** (R) | | 7 d (guest) · 24 h (app data) | **IMPLEMENTED / UNPROVEN** | every leg exists; the composed path has never been run. The destructive S5 drill is operator-gated and unrun (`06-offsite-connectivity.md:327`) |
|
||||||
| 10 | **Ransomware / malicious deletion inside the guest** | PBS offsite (the box **cannot** delete its own snapshots); **the restic repo is NOT protected the same way** | whole-guest restore from PBS to a point before the event | **operator** (SSH) | | 7 d | **PARTIAL** | R-89 proved the box is refused when deleting its own PBS snapshot (`CAMPAIGN-8…:514-515`). **R-95 (open, ranked #1):** the restic credential **can delete** — `readonly=False`, `forget --prune` runs from the box, and SFTP cannot express append-only (`OPEN-ITEMS.md:13`) **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open), and the depth that ships ON does not re-read pack contents (R-399).** |
|
| 10 | **Ransomware / malicious deletion inside the guest** | PBS offsite (the box **cannot** delete its own snapshots); **the restic repo is NOT protected the same way** | whole-guest restore from PBS to a point before the event | **operator** (SSH) | | 7 d | **PARTIAL** | R-89 proved the box is refused when deleting its own PBS snapshot (`CAMPAIGN-8…:514-515`). **R-95 (open, ranked #1):** the restic credential **can delete** — `readonly=False`, `forget --prune` runs from the box, and SFTP cannot express append-only (`OPEN-ITEMS.md:13`) **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. |
|
||||||
| 11 | **Hub lost** | every box's data plane, every tier, every Lane-1 route | none needed for recovery **of a box**; the hub itself restores from its Longhorn volume backup | operator (`kubectl`) | | 24 h + weekly (Longhorn `RETAIN 1` each) | **UNPROVEN** | a hub restore has never been performed. The backup target is `nfs://192.168.0.180` — **DooPlex itself** — and exactly **2** restore points exist (LIVE, INV Part D2.2) |
|
| 11 | **Hub lost** | every box's data plane, every tier, every Lane-1 route | none needed for recovery **of a box**; the hub itself restores from its Longhorn volume backup | operator (`kubectl`) | | 24 h + weekly (Longhorn `RETAIN 1` each) | **UNPROVEN** | a hub restore has never been performed. The backup target is `nfs://192.168.0.180` — **DooPlex itself** — and exactly **2** restore points exist (LIVE, INV Part D2.2) |
|
||||||
| 11b | *consequences while the hub is gone* | — | — | — | | | **[FACT]** | **day:** nothing customer-visible breaks; events queue (`settings.go:1466-1485`). **week:** the operator alarm plane is dark, no claim/reset codes, no config or floor convergence, no PBS-secret re-issue. **permanently:** escrow custody and break-glass credentials are gone (INV Part D2.4) |
|
| 11b | *consequences while the hub is gone* | — | — | — | | | **[FACT]** | **day:** nothing customer-visible breaks; events queue (`settings.go:1466-1485`). **week:** the operator alarm plane is dark, no claim/reset codes, no config or floor convergence, no PBS-secret re-issue. **permanently:** escrow custody and break-glass credentials are gone (INV Part D2.4) |
|
||||||
| 12 | **Offsite provider lost (Hetzner)** | everything on-premises: both drives, both whole-guest tiers | none needed — on-premises recovery is unaffected. Re-provision a new offsite target. | operator | | | **[FACT]** | **restic and PBS share the provider** (INV Part E.1). Whether they share an account and payment method is **UNKNOWN** → §11-D |
|
| 12 | **Offsite provider lost (Hetzner)** | everything on-premises: both drives, both whole-guest tiers | none needed — on-premises recovery is unaffected. Re-provision a new offsite target. | operator | | | **[FACT]** | **restic and PBS share the provider** (INV Part E.1). Whether they share an account and payment method is **UNKNOWN** → §11-D |
|
||||||
@@ -1079,7 +1079,7 @@ does **not** hold as written. → **R-108**
|
|||||||
| ~~R-359~~ | ~~The off-site restic store is never verified by anything, ever~~ | **CLOSED 2026-08-30, controller v0.227.0/v0.227.1.** A daily `offsite-integrity` job on **due-ness, not a weekday**; it takes the single-writer flag and SKIPS rather than waits (`resticStep` escalates to `unlock --remove-all` and is only safe while that flag is held). Three outcomes — skipped / unreachable / failed — because 'I could not look' is not 'I looked and it is broken'. **⚠ The depth that ships ON does NOT catch silent corruption:** measured, a pack corrupted without a size change returned `no errors were found`, exit 0; only `--read-data*` caught it. Choosing the depth is **R-399** |
|
| ~~R-359~~ | ~~The off-site restic store is never verified by anything, ever~~ | **CLOSED 2026-08-30, controller v0.227.0/v0.227.1.** A daily `offsite-integrity` job on **due-ness, not a weekday**; it takes the single-writer flag and SKIPS rather than waits (`resticStep` escalates to `unlock --remove-all` and is only safe while that flag is held). Three outcomes — skipped / unreachable / failed — because 'I could not look' is not 'I looked and it is broken'. **⚠ The depth that ships ON does NOT catch silent corruption:** measured, a pack corrupted without a size change returned `no errors were found`, exit 0; only `--read-data*` caught it. Choosing the depth is **R-399** |
|
||||||
| ~~R-397~~ | ~~`NotifyIntegrityOK`/`NotifyIntegrityFailed` had no caller and the product advertised a weekly check that did not exist~~ | **CLOSED 2026-08-30, controller v0.227.0.** Sixth built-but-never-wired instance: hub allowlist, Hungarian text, settings checkbox and debug button all existed; only the caller did not. `ok` is severity `info` and mails nobody by design |
|
| ~~R-397~~ | ~~`NotifyIntegrityOK`/`NotifyIntegrityFailed` had no caller and the product advertised a weekly check that did not exist~~ | **CLOSED 2026-08-30, controller v0.227.0.** Sixth built-but-never-wired instance: hub allowlist, Hungarian text, settings checkbox and debug button all existed; only the caller did not. `ok` is severity `info` and mails nobody by design |
|
||||||
| ~~R-399~~ | ~~The check reads the catalogue and never the data~~ | **CLOSED 2026-08-31, controller v0.228.0.** `monitoring.integrity.read_data_subset` now defaults to **`100%`**, so the weekly check downloads and re-hashes every stored byte. **The fact that made it necessary, and the sentence that should stop anyone turning it back down to save four seconds: the structure check PASSED a size-preserving pack corruption.** Measured on `demo-hp` 2026-08-30 — plain `restic check` reported `no errors were found` and exited 0 over a pack damaged without a size change; every read-data form caught it. Cost on that 134 MB store: 35.0 s structure vs 39.2 s at 100%. `off` (any case) returns a box to structure depth; an empty value means *not configured*, therefore the default; a malformed value WARNs and falls back to the DEFAULT, never to structure. A completed check over 5 minutes logs an operator WARN naming R-401 — **one data point, on one 134 MB store, so no rotation schedule, size threshold or bandwidth budget was invented from it.** Proven live at both depths 2026-08-31 with the restic argv observed from the guest |
|
| ~~R-399~~ | ~~The check reads the catalogue and never the data~~ | **CLOSED 2026-08-31, controller v0.228.0.** `monitoring.integrity.read_data_subset` now defaults to **`100%`**, so the weekly check downloads and re-hashes every stored byte. **The fact that made it necessary, and the sentence that should stop anyone turning it back down to save four seconds: the structure check PASSED a size-preserving pack corruption.** Measured on `demo-hp` 2026-08-30 — plain `restic check` reported `no errors were found` and exited 0 over a pack damaged without a size change; every read-data form caught it. Cost on that 134 MB store: 35.0 s structure vs 39.2 s at 100%. `off` (any case) returns a box to structure depth; an empty value means *not configured*, therefore the default; a malformed value WARNs and falls back to the DEFAULT, never to structure. A completed check over 5 minutes logs an operator WARN naming R-401 — **one data point, on one 134 MB store, so no rotation schedule, size threshold or bandwidth budget was invented from it.** Proven live at both depths 2026-08-31 with the restic argv observed from the guest |
|
||||||
| **R-87 (open) — AND IT IS NOT R-359** | The restic tier is never restore-TESTED | matrix row 4's route has no unattended proof. **Stated explicitly because the two rows sit next to each other and are easy to conflate: a `check` proves the STORE IS READABLE; a restore-test proves DATA COMES BACK OUT.** v0.227.0 did the first and nothing else. This row is untouched by it |
|
| **R-87 (open — SPIKED 2026-08-31, RE-SCOPE PROPOSED) — AND IT IS NOT R-359** | The restic tier is never restore-TESTED | **SPIKE VERDICT, `audits/SPIKE-restic-restore-test-2026-08-31.md`:** build the NARROW version, not the row as written. **Measured:** a scratch restore of all 8 apps costs **25 s / ≤213 MB scratch**, LESS than the 40.3 s weekly check beside it; but restic 0.14.0's `--verify` is a size-and-mtime reconciliation (a size-and-mtime-preserving corruption passed clean, red-proofed), `restic ls --json` carries no content hash, and the unit manifest hashes **4 918 B of a 213 231 242 B unit** — so **no reference for "correct" exists** (R-409). **Of the five drill-found restore defects R-353/354/356/358/403, an unattended scratch-restore would have caught ONE (R-356).** The value is elsewhere and the weekly check structurally cannot reach it: `check` proves the stored bytes are the stored bytes, never that we stored the RIGHT thing — a hollow unit backs up, checks and restores cleanly and recovers nothing (R-403, measured 2026-08-31). **Proposed re-scope, Viktor's call:** *prove the off-site snapshot still CONTAINS a recoverable unit* — one app a night, restored to scratch, checked against its own `manifest.json` via the existing `unitCarriesData`. Must use `--no-lock` and skip `unlockStale` (R-95's constraint is otherwise violated — R-407/R-408 record what the path writes today) and must take `acquireRunning`, which `RestoreOffboxScratch` does not. | matrix row 4's route has no unattended proof. **Stated explicitly because the two rows sit next to each other and are easy to conflate: a `check` proves the STORE IS READABLE; a restore-test proves DATA COMES BACK OUT.** v0.227.0 did the first and nothing else. This row is untouched by it |
|
||||||
|
|
||||||
### 10.3 Divergences that are documented elsewhere and are not re-opened here
|
### 10.3 Divergences that are documented elsewhere and are not re-opened here
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,465 @@
|
|||||||
|
# SPIKE — can the off-site (restic) copy be restore-tested without a person? (R-87)
|
||||||
|
|
||||||
|
**Date:** 2026-08-31 · **Venue:** `demo-hp` (Tier 0, disposable), guest 9201, controller **v0.230.0**
|
||||||
|
**Class:** spike. **No production code was written.** No version bump, no build, no deploy, no golden.
|
||||||
|
**Evidence:** `documentation/audits/evidence-spike-restic-restore-2026-08-31/` — 31 files, all pulled
|
||||||
|
off the box before teardown (R-320).
|
||||||
|
|
||||||
|
**Baselines re-confirmed at the start of the session:** `felhom-controller` `main` @
|
||||||
|
`2d802d75e88616d86cbade8a0e16965c2b85771c`, `v0.230.0`; `felhom.eu` @
|
||||||
|
`dddcc808be95d1c89b276b4d791491bad3c96bba` (clean tree, `HEAD == origin/main`); `felhom-agent`
|
||||||
|
`058b945`, `v0.130.0`.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## The one-paragraph answer
|
||||||
|
|
||||||
|
**restic 0.14.0 cannot tell us a restore produced correct files, and neither can anything else the
|
||||||
|
box holds today** — `--verify` exists but checks size and modification time, not content, and the
|
||||||
|
recovery unit's own manifest records a hash for 4 918 bytes of a 213 231 242-byte unit. **But an
|
||||||
|
unattended restore-test is far cheaper than expected:** restoring *every* app on the box — 8
|
||||||
|
snapshots, 774 MB logical — took **25 seconds**, against **40.3 seconds** for the weekly integrity
|
||||||
|
check sitting beside it. And the question worth asking is not the one R-87 was filed for. Of the five
|
||||||
|
restore-path defects human drills found between 2026-08-26 and 2026-08-31, an unattended
|
||||||
|
scratch-restore would have caught **one**. What it *would* catch, and what the weekly check
|
||||||
|
structurally cannot, is a snapshot that is perfectly intact and contains **nothing recoverable** —
|
||||||
|
the R-403 shape, measured live nine days ago. **Recommendation: build the narrow version (option C
|
||||||
|
below), and do not build the thing R-87 asks for.**
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Part 1 — the register correction (done, pushed as `6e550ae`)
|
||||||
|
|
||||||
|
**When and where it went wrong.** R-87 was moved into `CLOSED-ITEMS.md` by commit **`ef6ac6f`**,
|
||||||
|
2026-08-22, *"One register, enforced by a gate; closed work compressed into siblings
|
||||||
|
(R-376..R-378)"*. Established from `git log -S`, not inferred: that commit is the only one that ever
|
||||||
|
added an R-87 row to `CLOSED-ITEMS.md` and the only one that removed it from `OPEN-ITEMS.md`.
|
||||||
|
|
||||||
|
**It is a survivor of R-378, not a separate incident.** R-378 records that same sweep moving six
|
||||||
|
still-open rows — R-123, R-190, R-214, R-264, R-295, R-352 — and restoring them verbatim in the same
|
||||||
|
session. **R-87 is a seventh it missed.** Its state cell read
|
||||||
|
`READY — RE-RANKED UP 2026-08-03 (R-86 closed)`: the leading verdict is `READY`, and the word
|
||||||
|
`closed` later in the same cell describes a *different* row. Nine days in the wrong file, while
|
||||||
|
`OPEN-ITEMS.md`'s ranking paragraph ranked it fourth and pointed at nothing.
|
||||||
|
|
||||||
|
**The count, reproduced independently — and the predicate decides the answer.**
|
||||||
|
|
||||||
|
| predicate | rows convicted in `CLOSED-ITEMS.md` (of 151) | which |
|
||||||
|
|---|---|---|
|
||||||
|
| open word anywhere in the **row** | 144 | meaningless |
|
||||||
|
| open word anywhere in the **state cell** | 3 | R-87, **R-224**, **R-260** |
|
||||||
|
| open word in the **leading verdict** | **1** | R-87 |
|
||||||
|
|
||||||
|
R-224 and R-260 are genuinely closed; their long prose verdicts merely contain the words "open" and
|
||||||
|
"OPEN". **The task author's count of one is right, and it is right only under the leading-verdict
|
||||||
|
predicate** — which is R-378's own lesson, restated by measurement.
|
||||||
|
|
||||||
|
**The gate:** `scripts/closed_register_gate.py`, two rules — no open state word leading a
|
||||||
|
`CLOSED-ITEMS.md` row's verdict, and no `R-` id with a row in both registers. Red-proofed on both
|
||||||
|
rules; negative-controlled against the files **as they were pushed**, where it convicts R-87 by name
|
||||||
|
and exits 1. Registered as the 12th gate in `repo_gates.py` **after** it was green. Four residual
|
||||||
|
holes are named in its docstring.
|
||||||
|
|
||||||
|
**One thing the second rule turned up:** `R-398` also had a row in both registers — a deliberate
|
||||||
|
cross-reference stub. It is now prose beneath the table, not a table row.
|
||||||
|
|
||||||
|
**A duplicate this session did NOT fix:** `OPEN-ITEMS.md` carries two unrelated findings both
|
||||||
|
numbered **R-133** (`:267` hub `customer_id` uniqueness; `:273` plaintext break-glass credential).
|
||||||
|
Filed as **R-406**; deliberately not gated, because a within-register duplicate rule would fail on a
|
||||||
|
pre-existing row and a registered-but-failing gate refuses every push.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Q1 — what restic is actually running?
|
||||||
|
|
||||||
|
**Answer: restic 0.14.0, and every source comment asserting that is correct.**
|
||||||
|
|
||||||
|
From the running container on `demo-hp`, not from the Dockerfile:
|
||||||
|
|
||||||
|
```
|
||||||
|
restic 0.14.0 compiled with go1.19.8 on linux/amd64
|
||||||
|
/usr/bin/restic
|
||||||
|
/etc/debian_version → 12.15 (Debian bookworm, as the Dockerfile says)
|
||||||
|
```
|
||||||
|
|
||||||
|
Method: `pct exec 9201 -- docker exec felhom-controller restic version`, rc=0. Evidence
|
||||||
|
`02-q1-restic-version.txt`. The comments at `offbox_capture.go:15`, `offbox_restore.go:21`,
|
||||||
|
`offbox.go:726`, `offbox_progress.go:48,66` are **confirmed, not corrected**.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Q2 — what can that version verify about a RESTORE?
|
||||||
|
|
||||||
|
**Answer: `--verify` exists, and it does NOT verify content. It cannot tell us a restore produced
|
||||||
|
correct files.**
|
||||||
|
|
||||||
|
**`--verify` is real.** `restic restore --help` on the running container lists
|
||||||
|
`--verify verify restored files content`. Controls, because a `grep -c 0` must be earned:
|
||||||
|
|
||||||
|
| control | result |
|
||||||
|
|---|---|
|
||||||
|
| positive — `--target` present | 1 hit |
|
||||||
|
| `--verify` present | 1 hit, quoted verbatim above |
|
||||||
|
| negative — `--delete`, `--dry-run`, `--overwrite`, `--sparse` (all post-0.14) | 0 hits each |
|
||||||
|
| negative — `ZZZ-NOT-A-FLAG` | 0 hits |
|
||||||
|
|
||||||
|
Evidence `03-q2-restore-help.txt`, `04-q2-controls.txt`. **`--verify` and `--no-lock` both exist in
|
||||||
|
0.14.0 and neither appears anywhere in the controller source** — `grep -rn` over
|
||||||
|
`felhom-controller/controller/` returns rc=1 for both, with `--json`/`--target` as the positive
|
||||||
|
control (`05-q2-codebase-verify-grep.txt`).
|
||||||
|
|
||||||
|
**What `--verify` actually checks — measured, with a red-proof and a negative control.**
|
||||||
|
|
||||||
|
| test | what was done | result |
|
||||||
|
|---|---|---|
|
||||||
|
| cost | virgin restore of kimai (213 231 242 B, 7 files) with `--verify` | verify itself **131 ms**; restore+verify 3 438 ms |
|
||||||
|
| **red-proof** | corrupt one byte in a restored 160 MB tar, **size and mtime preserved**, re-run `restore --verify` | **PASSED clean, rc=0** — the corruption was not detected |
|
||||||
|
| control (size) | truncate the same file by 1 byte, re-run | restic silently **re-downloaded** it; verify reported "7 files, 133 ms", rc=0 |
|
||||||
|
| control (mtime) | corrupt content and bump mtime, re-run | restic silently **re-downloaded** it; rc=0 |
|
||||||
|
| negative control | same command against the untouched copy | identical output — so "passed" carries no information |
|
||||||
|
|
||||||
|
**131 ms cannot hash 213 MB.** Combined with the red-proof, `--verify` in 0.14.0 is a size-and-mtime
|
||||||
|
reconciliation that re-fetches anything that disagrees. It is useful — it makes a restore
|
||||||
|
self-repairing — and it is **not** a content check. Evidence `20-`, `21-`, `22-`.
|
||||||
|
|
||||||
|
**This is the point at which the spike's shape changed**, per §9's instruction to stop and
|
||||||
|
reconsider after Q1/Q2: the tool cannot supply the reference, so Q3 became the hard question.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Q3 — if not restic, then what is the reference for "correct"?
|
||||||
|
|
||||||
|
**Answer: there is none available to an unattended test today. The one hash record that exists covers
|
||||||
|
0.002 % of a unit's bytes.**
|
||||||
|
|
||||||
|
| candidate | verdict | why |
|
||||||
|
|---|---|---|
|
||||||
|
| restic `--verify` | **rejected** | size + mtime only — Q2's red-proof |
|
||||||
|
| the snapshot's own metadata | **rejected** | `restic ls --json` file nodes in 0.14.0 carry name, size, mode, uid/gid and three timestamps — **no content hash** (`24-q3-hash-coverage.txt`) |
|
||||||
|
| a hash the controller already records | **rejected as a content reference, kept as a completeness one** | see below |
|
||||||
|
| `restic check --read-data` | **already shipped, different question** | it proves the STORE's packs, never the restored files |
|
||||||
|
| a planted sentinel | **rejected** | a drill technique. An unattended test may not write data into a customer's app to have something to look for |
|
||||||
|
| the live data | **rejected** | it drifts by design; the snapshot is 12 h old by the time a check runs |
|
||||||
|
| restore twice and compare | **rejected** | proves determinism, not correctness |
|
||||||
|
|
||||||
|
**The hash record that exists, and its exact coverage.** The recovery unit's `manifest.json` carries
|
||||||
|
a `checksums` object. For kimai, measured on the restored unit:
|
||||||
|
|
||||||
|
```
|
||||||
|
checksums: .felhom.yml (2 235 B), app.yaml (488 B), docker-compose.yml (2 195 B) = 4 918 B
|
||||||
|
files in the unit: + kimai-mariadb.sql 48 217 B
|
||||||
|
+ kimai_kimai_db_data.tar 160 331 776 B
|
||||||
|
+ kimai_kimai_var.tar 52 845 056 B
|
||||||
|
+ manifest.json 1 275 B total 213 231 242 B
|
||||||
|
```
|
||||||
|
|
||||||
|
**4 918 of 213 231 242 bytes — 0.0023 %.** The three config files are hashed; the database dump and
|
||||||
|
the two volume tars, which are the recoverable data, are not. Filed as **R-409**.
|
||||||
|
|
||||||
|
**What the manifest CAN answer is a different and better question.** It declares `db_dumps` and
|
||||||
|
`volume_dumps` by name, and `unitCarriesData` (`r403_hollow.go:40`) already asks it. A restored unit
|
||||||
|
can therefore be checked for **completeness** — does every file the manifest declares exist — with
|
||||||
|
no new metadata, no new reference, and no content hash. That is the whole of the recommendation in
|
||||||
|
§Recommendation.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Q4 — what does one restore-test cost?
|
||||||
|
|
||||||
|
**Answer: about 4 seconds per app and 25 seconds for the whole box — cheaper than the weekly
|
||||||
|
integrity check it would sit beside. The cost is dominated by per-snapshot round-trip, not by data
|
||||||
|
volume.**
|
||||||
|
|
||||||
|
**Through the product's own path** (`POST /backup/offbox/restore` → `RestoreOffboxScratch`), timed
|
||||||
|
from the POST to the completion log line:
|
||||||
|
|
||||||
|
| app | mode | logical size | wall clock |
|
||||||
|
|---|---|---|---|
|
||||||
|
| docmost | unit | 118 207 270 B | **9 s** (13:40:10 → 13:40:19) |
|
||||||
|
| kimai | full | 213 231 242 B | **11 s** (13:42:54 → 13:43:05) |
|
||||||
|
| opengist | unit | 185 664 B | ~8 s |
|
||||||
|
|
||||||
|
**Raw restic, all eight snapshots back to back** (`26-q4-all-apps-cost.txt`):
|
||||||
|
|
||||||
|
| app | logical | restored to disk | ms |
|
||||||
|
|---|---|---|---|
|
||||||
|
| privatebin | 2 123 006 | 2 159 870 | 2 628 |
|
||||||
|
| opengist | 185 664 | 222 528 | 2 253 |
|
||||||
|
| calibre-web | 5 808 703 | 5 878 335 | 3 642 |
|
||||||
|
| paperless-ngx | 83 473 654 | 83 547 382 | 3 062 |
|
||||||
|
| bookstack | 166 641 768 | 166 682 728 | 2 745 |
|
||||||
|
| docmost | 118 207 270 | 118 248 230 | 3 676 |
|
||||||
|
| romm | 184 706 816 | 184 747 776 | 3 978 |
|
||||||
|
| kimai | 213 231 242 | 213 272 202 | 3 202 |
|
||||||
|
| **all eight** | **774 378 123** | **774 759 051** | **25 s** |
|
||||||
|
|
||||||
|
**185 KB takes 2.25 s and 213 MB takes 3.20 s.** Nearly all of it is fixed per-snapshot cost —
|
||||||
|
opening the repo, loading the index, one SFTP session. Data adds roughly **1 s per 200 MB**
|
||||||
|
(≈ 67 MB/s on this link).
|
||||||
|
|
||||||
|
**The cache is not hiding the cost.** `/root/.cache/restic` is **1.1 MB** — index and metadata only.
|
||||||
|
Cached 3 198 ms vs `--no-cache` 5 423 ms for kimai; the trees are byte-identical (`diff -r` → YES).
|
||||||
|
Pack data always crosses the wire. Evidence `25-q4-cache-effect.txt`.
|
||||||
|
|
||||||
|
**Peak scratch:** the restore writes the **full logical size** — 213 272 202 B for the largest app.
|
||||||
|
Sequential-with-cleanup needs only the largest app; all-at-once needs 774 MB.
|
||||||
|
|
||||||
|
**Against R-359's numbers.** R-359 measured 35.0 s at structure depth and 39.2 s at 100 % read-data
|
||||||
|
on a 140 829 678 B store. Re-measured today through the product's own debug button on a
|
||||||
|
141 959 062 B store: **40 257 ms at 100 %** (`15-q5-integrity-run.txt`). So:
|
||||||
|
|
||||||
|
> **restore-testing every app on the box (25 s) costs LESS than one weekly integrity check (40 s).**
|
||||||
|
> Same order of magnitude, and on the cheaper side of it.
|
||||||
|
|
||||||
|
**Extrapolation — labelled as extrapolation.** R-401 records that the check's cost tracks the index
|
||||||
|
and a read-data run's tracks the data. **A restore tracks the data too, plus a fixed per-snapshot
|
||||||
|
cost.** With 8 apps and the measured ≈ 67 MB/s:
|
||||||
|
|
||||||
|
| store | fixed (8 × ~2 s) | data | total |
|
||||||
|
|---|---|---|---|
|
||||||
|
| today, 774 MB logical | 16 s | ~9 s | **25 s (measured)** |
|
||||||
|
| 10× — 7.7 GB | 16 s | ~115 s | ~2.2 min |
|
||||||
|
| 100× — 77 GB | 16 s | ~19 min | ~19 min |
|
||||||
|
|
||||||
|
**Why it may not hold.** One store, one link, one afternoon. This is DooPlex→Hetzner; a customer's
|
||||||
|
domestic line is the real variable, and restore is the *download* direction, usually the faster one
|
||||||
|
on such a line. Scratch space becomes the binding constraint long before time does: at 100× the
|
||||||
|
largest app would need ~21 GB of free scratch, and `RestoreOffboxScratch`'s headroom gate would start
|
||||||
|
refusing. **Unknown, and what would settle it:** a measurement on a store above 10 GB. None exists.
|
||||||
|
This is the same single-data-point hole R-401 already owns.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Q5 — contention
|
||||||
|
|
||||||
|
**Answer: skip-if-busy stays right, because the operation is seconds and not minutes — but the
|
||||||
|
restore path takes NO single-writer flag at all today, which is a bigger problem than contention.**
|
||||||
|
|
||||||
|
**The configured window on this box**, read from the scheduler's own registrations rather than
|
||||||
|
assumed (`28-q5-schedule.txt`, times are **CEST** — the guest scheduler runs in local time, not UTC):
|
||||||
|
|
||||||
|
| job | time | measured duration |
|
||||||
|
|---|---|---|
|
||||||
|
| db-dump | 02:30 | — |
|
||||||
|
| tier2-backup | 03:30 | — |
|
||||||
|
| **offbox-backup** | **04:15** | **2m52s** (last run, `last_duration`) |
|
||||||
|
| offsite-abandon-sweep | 05:10 | — |
|
||||||
|
| **offsite-integrity** | **06:00** | **40.3 s** at 100 % depth |
|
||||||
|
|
||||||
|
A restore-test of all eight apps holds anything it holds for **25 s** — one seventh of the nightly
|
||||||
|
off-site backup, and shorter than the check beside it. There is an empty gap from ~04:18 to 06:00.
|
||||||
|
**Skip-if-busy remains the right policy**, and it is not the "minutes rather than seconds" case the
|
||||||
|
task worried about. The `integrityCheckTimeout` reasoning transfers unchanged.
|
||||||
|
|
||||||
|
**The real finding here.** `offbox_integrity.go:28` states the invariant:
|
||||||
|
|
||||||
|
> *"Every off-site operation takes `acquireRunning` for exactly that reason."*
|
||||||
|
|
||||||
|
**It does not.** `grep -rn 'acquireRunning()'` finds nine non-test callers; `RestoreOffboxScratch`
|
||||||
|
(`offbox_restore.go:234`) is **not** among them. `restore_wizard.go:174` says the same thing
|
||||||
|
independently — *"`RestoreOffboxScratch` never acquires it at all"* — and the UI works around it with
|
||||||
|
a separate display flag. So a scheduled, unattended caller built on `RestoreOffboxScratch` today
|
||||||
|
would run with **no single-writer flag**, which is precisely the hazard the integrity file's header
|
||||||
|
is shaped around. **A comment asserting an invariant with no test pinning it.** Filed as **R-408**.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Q6 — what does a restore-test WRITE to the repository?
|
||||||
|
|
||||||
|
**Answer: `restic restore` writes nothing — it does not even take a lock. The product's restore path
|
||||||
|
writes anyway, because it runs `restic unlock` before every restore. And `restic check` — the thing
|
||||||
|
whose comment says it never writes — DOES take a lock.**
|
||||||
|
|
||||||
|
**Method.** Two observers inside the container, both proven before being believed:
|
||||||
|
|
||||||
|
1. a lock sampler running `restic list locks --no-lock` every ~4 s (the observer itself never writes);
|
||||||
|
2. an argv sampler reading `/proc/*/cmdline` every 0.2 s, with the repo URL and `sftp.command`
|
||||||
|
redacted at the source.
|
||||||
|
|
||||||
|
**Positive control for the lock sampler — it works.** The product's own integrity check ran
|
||||||
|
13:41:28 → 13:42:11:
|
||||||
|
|
||||||
|
```
|
||||||
|
13:41:30 locks=0
|
||||||
|
13:41:34 locks=1 ids=81fd4d4200d848466e18cb7a9d8e0d4300d95707c68417449178c3023531eecf
|
||||||
|
... nine consecutive samples, one lock ...
|
||||||
|
13:42:10 locks=1 ids=81fd…
|
||||||
|
13:42:14 locks=0
|
||||||
|
```
|
||||||
|
|
||||||
|
**So `restic check` writes a lock file to the repository.** `offbox_integrity.go:255` says
|
||||||
|
*"It NEVER writes to the repository: `check` is a read verb, and nothing here prunes, forgets,
|
||||||
|
unlocks or backs up."* The three named verbs are correct; the sentence's headline is not. Filed as
|
||||||
|
**R-407**.
|
||||||
|
|
||||||
|
**The restore, with the same proven instrument:** locks=0 across every sample inside both restore
|
||||||
|
windows (13:40:11/:15/:19 for docmost, 13:42:57/:43:01/:05 for kimai). **restic 0.14.0's `restore`
|
||||||
|
does not lock the repository.**
|
||||||
|
|
||||||
|
**What the product runs, observed argv (redacted), one restore of `opengist`:**
|
||||||
|
|
||||||
|
```
|
||||||
|
13:44:41 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> snapshots latest --tag opengist --json
|
||||||
|
13:44:44 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> unlock
|
||||||
|
13:44:46 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> restore b5aa8f9b --target … --include …
|
||||||
|
```
|
||||||
|
|
||||||
|
Three invocations. The middle one is `unlockStale` (`offbox_restore.go:289` → `offbox.go:743`), which
|
||||||
|
runs **unconditionally before every restore** and is a **delete verb against `locks/`**. With no
|
||||||
|
stale lock present it deletes nothing — but it is a write-capable command on the exact path R-87
|
||||||
|
wants to run unattended. `resticStep`'s escalation to `unlock --remove-all` (`offbox.go:760-775`)
|
||||||
|
fires only on `repository is already locked`, which a restore cannot provoke by itself now that we
|
||||||
|
know restore takes no lock.
|
||||||
|
|
||||||
|
**Against §5's constraint** — *"R-95 still applies: that credential can delete, so a restic
|
||||||
|
restore-test must never be able to write to the repo"*:
|
||||||
|
|
||||||
|
- **The lead in the task was right in direction and wrong in mechanism.** The write is not
|
||||||
|
`resticStep`'s escalation; it is `unlockStale`, one line earlier and unconditional.
|
||||||
|
- **The constraint IS satisfiable, cheaply, and both mechanisms already exist in 0.14.0 and are
|
||||||
|
unused:** a restore-test that passes `--no-lock` and skips `unlockStale` writes nothing to the
|
||||||
|
repository at all. That is a design note for whoever builds it — **this spike did not fix it**, per
|
||||||
|
§7.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Q7 — what would an unattended restore-test catch that the weekly 100 % check does not?
|
||||||
|
|
||||||
|
**Answer: of the five defects human drills found in the last six days, one. But that is the wrong
|
||||||
|
scoreboard, and the right one has a much better answer.**
|
||||||
|
|
||||||
|
| defect | would an unattended scratch-restore have caught it? | how / why not |
|
||||||
|
|---|---|---|
|
||||||
|
| **R-353** — a **local** unit restore reported a bare completion whether it returned a dataset or nothing | **NO** | different code path entirely (`restore_unit.go`). An off-site restore-test never enters it |
|
||||||
|
| **R-354** — the off-site full restore has no named-volume replay leg: the tar reaches the scratch and is never replayed | **NO** | the defect is *after* the scratch. A test that stops at the scratch sees a correct scratch. Going further means replaying into a live app, which an unattended test must not do |
|
||||||
|
| **R-356** — the off-site restore refused every app with no data drive (40 of 53 catalogue templates) | **YES** | the test calls `RestoreOffboxScratch`, gets a refusal, and `rerr != nil`. On this box five of eight apps live on `/mnt/sys_drive` with no drive — it would have fired on the first night |
|
||||||
|
| **R-358** — `OffboxFullScratchReady` asked "non-empty directory", which is what a *failed* restic run leaves | **NO** | needs a restore that fails part-way. On a healthy run the broken and the fixed predicate agree |
|
||||||
|
| **R-403** — a hollow unit mirrored over a complete one with `--delete`; 120 082 104 B → 7 036 B, recorded as success | **NO as filed** — but **YES for the shape** | R-403 destroyed a *local* second-drive copy. What a restore-test sees is the consequence: once a hollow unit is captured off-site, the snapshot is perfectly intact and contains nothing recoverable |
|
||||||
|
|
||||||
|
**One of five. If the question is "does our restore code work", the honest answer is: the drills
|
||||||
|
already answer it, they answer it better, and automating a worse version of it is not worth an
|
||||||
|
evening.**
|
||||||
|
|
||||||
|
**The scoreboard that matters is different, and the weekly check structurally cannot play on it.**
|
||||||
|
|
||||||
|
> `restic check --read-data-subset=100%` proves that **the bytes we stored are the bytes we stored.**
|
||||||
|
> It cannot tell us **we stored the wrong thing.**
|
||||||
|
|
||||||
|
A hollow recovery unit — no database dump, no volume tar — backs up cleanly, checks cleanly at 100 %
|
||||||
|
depth, restores cleanly, and recovers nothing. **R-403 proved that shape is real, on this fleet, nine
|
||||||
|
days ago, measured in bytes.** Nothing in the product asks the question today, on any tier, at any
|
||||||
|
cadence. A restore-test is simply the cheapest place to ask it, because the manifest that answers it
|
||||||
|
travels inside the snapshot.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Recommendation to Viktor — three options, with costs
|
||||||
|
|
||||||
|
**Option A — do not build it.**
|
||||||
|
*Cost:* nothing. *What you get:* the weekly 100 % check keeps proving the stored bytes; the restore
|
||||||
|
code keeps being proven by your drills.
|
||||||
|
*What you lose:* nothing that has bitten yet — and the R-403 shape stays invisible until a customer
|
||||||
|
needs the data. **This is a defensible answer** and it is the one R-87's original framing deserves.
|
||||||
|
|
||||||
|
**Option B — a scheduled attended drill instead.**
|
||||||
|
*Cost:* one of your evenings, monthly. *What you get:* everything a person can see, including the
|
||||||
|
R-354 class that stops at the scratch.
|
||||||
|
*Honest objection:* this is what already happens, and it is what found all five defects. Scheduling
|
||||||
|
it changes nothing except that it now has a date. **Low value for the price.**
|
||||||
|
|
||||||
|
**Option C — build the NARROW unattended test: one app per night, rotating, restored to scratch, and
|
||||||
|
checked against its own manifest. RECOMMENDED.**
|
||||||
|
|
||||||
|
*What it does, in one sentence:* restore the newest off-site snapshot of one app into the throwaway
|
||||||
|
scratch, assert that every file the unit's `manifest.json` declares is present, record which snapshot
|
||||||
|
was proved, delete the scratch.
|
||||||
|
|
||||||
|
*Measured cost, not estimated:* **~4 s and ≤ 213 MB of scratch per night** (one app), or 25 s for all
|
||||||
|
eight. Off-site traffic: one restore's worth, ≈ the app's size. **Less than the weekly integrity
|
||||||
|
check already running beside it.** No new metadata, no new reference, no content hashes — it reuses
|
||||||
|
`unitCarriesData`'s existing manifest read.
|
||||||
|
|
||||||
|
*What it catches:* R-356 outright, and the R-403 class — a snapshot that is intact and empty — which
|
||||||
|
nothing else in the product asks about.
|
||||||
|
*What it does not catch, stated so nobody expects it to:* R-353, R-354, R-358. Those stay drill work.
|
||||||
|
|
||||||
|
*Three things it must be built with, all established by this spike:*
|
||||||
|
1. **`--no-lock`, and skip `unlockStale`** — then it writes nothing to the repository and R-95's
|
||||||
|
constraint is honoured for real (Q6).
|
||||||
|
2. **Take `acquireRunning`** — `RestoreOffboxScratch` does not, and the whole off-site
|
||||||
|
single-writer story assumes every operation does (Q5, R-408).
|
||||||
|
3. **Record the SNAPSHOT it proved, not a timestamp** — R-87's own row already says this, and R-86
|
||||||
|
built the per-archive due-ness model to copy.
|
||||||
|
|
||||||
|
**If you do nothing:** the weekly check keeps running and keeps being right about the bytes. The
|
||||||
|
first time a hollow unit reaches the off-site store, nothing will notice, and the discovery will be a
|
||||||
|
customer's restore. That is not a hypothetical shape — it is R-403, measured on 2026-08-31.
|
||||||
|
|
||||||
|
**I would pick C**, scoped exactly as above. It is cheaper than the check beside it, it asks a
|
||||||
|
question nothing else asks, and it needs no invention.
|
||||||
|
|
||||||
|
**R-87 itself should be RE-SCOPED, not built as written** — from *"restore-test the restic tier"* to
|
||||||
|
*"prove the off-site snapshot still contains a recoverable unit"*. That is your call, so R-87 stays
|
||||||
|
open carrying this verdict.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Register rows opened by this spike
|
||||||
|
|
||||||
|
| id | what |
|
||||||
|
|---|---|
|
||||||
|
| **R-405** | R-87 was mis-filed by `ef6ac6f`; corrected + gated (CLOSED same session) |
|
||||||
|
| **R-406** | two unrelated findings share the id R-133 in `OPEN-ITEMS.md` |
|
||||||
|
| **R-407** | `restic check` DOES take a repository lock; `offbox_integrity.go:255` says it never writes |
|
||||||
|
| **R-408** | `RestoreOffboxScratch` takes no `acquireRunning`; `offbox_integrity.go:28` asserts every off-site operation does |
|
||||||
|
| **R-409** | the recovery-unit manifest hashes 4 918 B of a 213 231 242 B unit — the data files have no recorded hash |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Probes, teardown and side-effects
|
||||||
|
|
||||||
|
**All three layers, and two of them really are "nothing was left".**
|
||||||
|
|
||||||
|
| layer | created | removed |
|
||||||
|
|---|---|---|
|
||||||
|
| PVE host `demo-hp` `/root` | 6 helper scripts + one 0600 password file | all removed; `ls \| grep` returns nothing |
|
||||||
|
| guest 9201 `/root`, `/tmp` | 9 helper/session files + `spike-env.sh` | all removed; grep returns nothing |
|
||||||
|
| container `/tmp` | `spike-env.sh`, two samplers, two logs, two run-flags | all removed; `/tmp` lists empty |
|
||||||
|
|
||||||
|
**Scratch directories:** four (`docmost`, `kimai`, `privatebin`, `opengist`) were created by the
|
||||||
|
restores this spike drove and were **removed**. Three (`bookstack`, `calibre-web`, `paperless-ngx`)
|
||||||
|
pre-date this session and were **left alone** — this spike never restored them.
|
||||||
|
|
||||||
|
**Two deliberate state changes on the box, recorded rather than hidden:**
|
||||||
|
|
||||||
|
1. The integrity check I ran to control the lock observer **recorded its verdict** — the box's
|
||||||
|
`last_integrity_check` moved to `2026-08-31T13:41:28Z`, `ok=true`, and
|
||||||
|
`last_integrity_depth` moved from `structure` to **`100%`**. Due-ness advanced by seven days. That
|
||||||
|
is the product behaving correctly; it is not a repair and it is not damage.
|
||||||
|
2. Two logins and four restores appear in the controller log and in the customer-visible restore-op
|
||||||
|
history.
|
||||||
|
|
||||||
|
**Nothing was written to the off-site repository by hand.** No prune, no forget, no `unlock` issued
|
||||||
|
by me. Every write observed in Q6 was the product's own.
|
||||||
|
|
||||||
|
**No controller code changed. No golden is owed.** The fleet stays on v0.230.0. The one delivery debt
|
||||||
|
that exists — `golden_currency_gate.py` red because v0.230.0 has no golden — was **already red at
|
||||||
|
`dddcc80`** before this session began, and belongs to the v0.230.0 release, not to this task.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Observations noticed and not acted on
|
||||||
|
|
||||||
|
- **`paperless-ngx` and `filebrowser` have no off-site snapshot** under the tags the inventory uses —
|
||||||
|
`paperless` returns nothing and `paperless-ngx` returns one; `filebrowser` returns none at all.
|
||||||
|
`filebrowser` is infrastructure, so that may be correct. Not chased.
|
||||||
|
- **`CLOSED-ITEMS.md` has two malformed rows** — R-399 and R-400 supply two columns where the table
|
||||||
|
declares four, so they render with no `Shipped` and no `Evidence`. The new gate prints them as a
|
||||||
|
warning rather than convicting, because an empty state cell is not an open state word.
|
||||||
|
- **Two rows carry a `|` inside their body** (R-309, R-351), which shifts their own cells. Named as
|
||||||
|
the new gate's first residual hole.
|
||||||
|
- **The controller image has no `ps` and no `python3`** — worth knowing before writing any probe that
|
||||||
|
runs inside it. `/proc/*/cmdline` is the substitute that works.
|
||||||
|
- **`git push --no-verify` was used for this session's Part 1 commit** — bypass **#7**, for the reason
|
||||||
|
R-404 exists: a documentation-only push met `golden_currency_gate.py`. Recorded here because R-404
|
||||||
|
counts them.
|
||||||
@@ -0,0 +1,4 @@
|
|||||||
|
demo-hp
|
||||||
|
15:32:56 up 9 days, 21:48, 1 user, load average: 0.70, 0.63, 0.57
|
||||||
|
VMID Status Lock Name
|
||||||
|
9201 running demo-hp
|
||||||
@@ -0,0 +1,20 @@
|
|||||||
|
perl: warning: Setting locale failed.
|
||||||
|
perl: warning: Please check that your locale settings:
|
||||||
|
LANGUAGE = (unset),
|
||||||
|
LC_ALL = (unset),
|
||||||
|
LC_CTYPE = "UTF-8",
|
||||||
|
LC_NUMERIC = (unset),
|
||||||
|
LC_COLLATE = (unset),
|
||||||
|
LC_TIME = (unset),
|
||||||
|
LC_MESSAGES = (unset),
|
||||||
|
LC_MONETARY = (unset),
|
||||||
|
LC_ADDRESS = (unset),
|
||||||
|
LC_IDENTIFICATION = (unset),
|
||||||
|
LC_MEASUREMENT = (unset),
|
||||||
|
LC_PAPER = (unset),
|
||||||
|
LC_TELEPHONE = (unset),
|
||||||
|
LC_NAME = (unset),
|
||||||
|
LANG = "en_US.UTF-8"
|
||||||
|
are supported and installed on your system.
|
||||||
|
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||||
|
felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.230.0 Up 59 minutes (healthy)
|
||||||
@@ -0,0 +1,4 @@
|
|||||||
|
restic 0.14.0 compiled with go1.19.8 on linux/amd64
|
||||||
|
restic-rc=0
|
||||||
|
/usr/bin/restic
|
||||||
|
12.15
|
||||||
@@ -0,0 +1,64 @@
|
|||||||
|
LANGUAGE = (unset),
|
||||||
|
LC_ALL = (unset),
|
||||||
|
LC_CTYPE = "UTF-8",
|
||||||
|
LC_NUMERIC = (unset),
|
||||||
|
LC_COLLATE = (unset),
|
||||||
|
LC_TIME = (unset),
|
||||||
|
LC_MESSAGES = (unset),
|
||||||
|
LC_MONETARY = (unset),
|
||||||
|
LC_ADDRESS = (unset),
|
||||||
|
LC_IDENTIFICATION = (unset),
|
||||||
|
LC_MEASUREMENT = (unset),
|
||||||
|
LC_PAPER = (unset),
|
||||||
|
LC_TELEPHONE = (unset),
|
||||||
|
LC_NAME = (unset),
|
||||||
|
LANG = "en_US.UTF-8"
|
||||||
|
|
||||||
|
The "restore" command extracts the data from a snapshot from the repository to
|
||||||
|
a directory.
|
||||||
|
|
||||||
|
The special snapshot "latest" can be used to restore the latest snapshot in the
|
||||||
|
repository.
|
||||||
|
|
||||||
|
EXIT STATUS
|
||||||
|
===========
|
||||||
|
|
||||||
|
Exit status is 0 if the command was successful, and non-zero if there was any error.
|
||||||
|
|
||||||
|
Usage:
|
||||||
|
restic restore [flags] snapshotID
|
||||||
|
|
||||||
|
Flags:
|
||||||
|
-e, --exclude pattern exclude a pattern (can be specified multiple times)
|
||||||
|
-h, --help help for restore
|
||||||
|
-H, --host stringArray only consider snapshots for this host when the snapshot ID is "latest" (can be specified multiple times)
|
||||||
|
--iexclude --exclude same as --exclude but ignores the casing of filenames
|
||||||
|
--iinclude --include same as --include but ignores the casing of filenames
|
||||||
|
-i, --include pattern include a pattern, exclude everything else (can be specified multiple times)
|
||||||
|
--path path only consider snapshots which include this (absolute) path for snapshot ID "latest"
|
||||||
|
--tag taglist only consider snapshots which include this taglist for snapshot ID "latest" (default [])
|
||||||
|
-t, --target string directory to extract data to
|
||||||
|
--verify verify restored files content
|
||||||
|
|
||||||
|
Global Flags:
|
||||||
|
--cacert file file to load root certificates from (default: use system certificates)
|
||||||
|
--cache-dir directory set the cache directory. (default: use system default cache directory)
|
||||||
|
--cleanup-cache auto remove old cache directories
|
||||||
|
--compression mode compression mode (only available for repository format version 2), one of (auto|off|max) (default auto)
|
||||||
|
--insecure-tls skip TLS certificate verification when connecting to the repository (insecure)
|
||||||
|
--json set output mode to JSON for commands that support it
|
||||||
|
--key-hint key key ID of key to try decrypting first (default: $RESTIC_KEY_HINT)
|
||||||
|
--limit-download int limits downloads to a maximum rate in KiB/s. (default: unlimited)
|
||||||
|
--limit-upload int limits uploads to a maximum rate in KiB/s. (default: unlimited)
|
||||||
|
--no-cache do not use a local cache
|
||||||
|
--no-lock do not lock the repository, this allows some operations on read-only repositories
|
||||||
|
-o, --option key=value set extended option (key=value, can be specified multiple times)
|
||||||
|
--pack-size uint set target pack size in MiB, created pack files may be larger (default: $RESTIC_PACK_SIZE)
|
||||||
|
--password-command command shell command to obtain the repository password from (default: $RESTIC_PASSWORD_COMMAND)
|
||||||
|
-p, --password-file file file to read the repository password from (default: $RESTIC_PASSWORD_FILE)
|
||||||
|
-q, --quiet do not output comprehensive progress report
|
||||||
|
-r, --repo repository repository to backup to or restore from (default: $RESTIC_REPOSITORY)
|
||||||
|
--repository-file file file to read the repository location from (default: $RESTIC_REPOSITORY_FILE)
|
||||||
|
--tls-client-cert file path to a file containing PEM encoded TLS client certificate and private key
|
||||||
|
-v, --verbose n be verbose (specify multiple times or a level using --verbose=n, max level/times is 3)
|
||||||
|
=== rc=0 ===
|
||||||
@@ -0,0 +1,11 @@
|
|||||||
|
--- POSITIVE CONTROL: a flag that MUST be there ---
|
||||||
|
1
|
||||||
|
--- the verify flag, verbatim ---
|
||||||
|
--verify verify restored files content
|
||||||
|
--- NEGATIVE CONTROL: flags added AFTER 0.14, must be absent ---
|
||||||
|
--delete hits=0
|
||||||
|
--dry-run hits=0
|
||||||
|
--overwrite hits=0
|
||||||
|
--sparse hits=0
|
||||||
|
--- NEGATIVE CONTROL 2: a string that cannot be there ---
|
||||||
|
0
|
||||||
+8
@@ -0,0 +1,8 @@
|
|||||||
|
=== '--verify' anywhere in controller/ (all files, incl gitignored cmd/) ===
|
||||||
|
grep-rc=1
|
||||||
|
|
||||||
|
=== POSITIVE CONTROL: '--no-lock' ===
|
||||||
|
grep-rc=1
|
||||||
|
|
||||||
|
=== every restic flag token used in internal/backup/ (non-test) ===
|
||||||
|
"--delete" "--group-by" "--ignore-existing" "--include" "--itemize-changes" "--json" "--keep-daily" "--keep-monthly" "--keep-weekly" "--mode" "--prune" "--read-data" "--remove-all" "--rm" "--tag" "--target"
|
||||||
@@ -0,0 +1,48 @@
|
|||||||
|
felhom-controller Up About an hour (healthy)
|
||||||
|
docmost Up About an hour (healthy)
|
||||||
|
docmost-redis Up About an hour (healthy)
|
||||||
|
docmost-postgres Up About an hour (healthy)
|
||||||
|
romm Up 4 hours (healthy)
|
||||||
|
romm-db Up 4 hours (healthy)
|
||||||
|
romm-redis Up 4 hours (healthy)
|
||||||
|
privatebin Up 4 hours (healthy)
|
||||||
|
paperless-webserver Up 4 hours (healthy)
|
||||||
|
paperless-postgres Up 4 hours (healthy)
|
||||||
|
paperless-redis Up 4 hours (healthy)
|
||||||
|
opengist Up 4 hours (healthy)
|
||||||
|
kimai Up 4 hours (healthy)
|
||||||
|
kimai-db Up 4 hours (healthy)
|
||||||
|
calibre-web Up 4 hours (healthy)
|
||||||
|
bookstack Up 4 hours (healthy)
|
||||||
|
bookstack-db Up 4 hours (healthy)
|
||||||
|
filebrowser Up 9 days (healthy)
|
||||||
|
cloudflared Up 9 days
|
||||||
|
traefik Up 9 days
|
||||||
|
=== guest disk ===
|
||||||
|
Filesystem Size Used Avail Use% Mounted on
|
||||||
|
/dev/mapper/pve-vm--9201--disk--0 32G 957M 29G 4% /
|
||||||
|
/dev/mapper/pve-vm--9201--disk--1 69G 12G 55G 18% /var/lib/felhom
|
||||||
|
/dev/mapper/pve-root 39G 17G 21G 46% /mnt/felhom-drives
|
||||||
|
none 492K 4.0K 488K 1% /dev
|
||||||
|
udev 14G 0 14G 0% /dev/tty
|
||||||
|
overlay 69G 12G 55G 18% /var/lib/docker/overlay2/0e82b836f4879d75df7380c71546ade764fc20c0f1c7e1d7905ab790ddbc7a4e/merged
|
||||||
|
overlay 69G 12G 55G 18% /var/lib/docker/overlay2/1924a292e532a1948a03c280bce4b21f48f109765d427ca110044f5e1469fcfe/merged
|
||||||
|
overlay 69G 12G 55G 18% /var/lib/docker/overlay2/ba706ad1711d7d7d8b31250f124caac3e889a66be994e9bba959d7e9542ee76a/merged
|
||||||
|
/dev/nvme0n1 938G 6.0G 885G 1% /mnt/felhom-drives/hdd_1
|
||||||
|
overlay 69G 12G 55G 18% /var/lib/docker/overlay2/c52ca60013e4af6d6d93a4ca471d9c1dffc7f99a9923869e4d72f47e17dfa44f/merged
|
||||||
|
overlay 69G 12G 55G 18% /var/lib/docker/overlay2/e79e25da71576531d24348a67d88e2c7d1c9e4105622ae95cd7a40bd6c52cb57/merged
|
||||||
|
overlay 69G 12G 55G 18% /var/lib/docker/overlay2/8131f6f54062f0094af0e03a1ad26a7c77ed6991e27a10211403cfbb7cb479a1/merged
|
||||||
|
overlay 69G 12G 55G 18% /var/lib/docker/overlay2/932ec3a6439333b5ecf8b39cb68795e15bdd1dc4a7fb831cae46bd9625ebaeb2/merged
|
||||||
|
overlay 69G 12G 55G 18% /var/lib/docker/overlay2/a2d41f6c1de3963c93bce576a41d0922695ec4d1460b4ebdd9a398245028182e/merged
|
||||||
|
overlay 69G 12G 55G 18% /var/lib/docker/overlay2/c0efcc2fbbb80d78ee5d764a3f9f6ab9e622d1df2faac473068aa0486f2c3d66/merged
|
||||||
|
overlay 69G 12G 55G 18% /var/lib/docker/overlay2/b976192342358497830be3cd329d63ca9546550da2f849b2baa44c041ded1c17/merged
|
||||||
|
overlay 69G 12G 55G 18% /var/lib/docker/overlay2/a3d1768da870b1d86015d6850275ebba59cb905e4415f9cd399c1bf15b6d08e1/merged
|
||||||
|
overlay 69G 12G 55G 18% /var/lib/docker/overlay2/a0f134a5b31237fbc56ea1fb38113d3a0352f9d4a3d36e5e40b3cff874b42366/merged
|
||||||
|
overlay 69G 12G 55G 18% /var/lib/docker/overlay2/1400ead8a09768fd52c279ea50cffbd5b87d061ef83fb93a0fa0685d7e0b027b/merged
|
||||||
|
overlay 69G 12G 55G 18% /var/lib/docker/overlay2/3723f8b898e197ad12b8fcaef66c08569cc9040d1efd1cba7d8797b5a7f0b14f/merged
|
||||||
|
overlay 69G 12G 55G 18% /var/lib/docker/overlay2/25ee94f6cfa646a342dd751ca14f58e56f3227aaaf9160d26409c4cc23994453/merged
|
||||||
|
overlay 69G 12G 55G 18% /var/lib/docker/overlay2/9e92315ab21675a0f271efb4c8eb0ed792a4de69dd58ec5e51300dfa3852c4b9/merged
|
||||||
|
overlay 69G 12G 55G 18% /var/lib/docker/overlay2/bb1ad0c258566bc3d4013171500fa6a7ed30a80ee620bc93ef8adb939eb0c166/merged
|
||||||
|
overlay 69G 12G 55G 18% /var/lib/docker/overlay2/2461db88cd4e3b967ab86cdc6a41de3f62b433ccbc1ca0424876a79ac0ba8d10/merged
|
||||||
|
overlay 69G 12G 55G 18% /var/lib/docker/overlay2/ba1fcae5ed394c9dbdee666392cc61400ef700a69f30ad98fe2de38ef43d7a24/merged
|
||||||
|
overlay 69G 12G 55G 18% /var/lib/docker/overlay2/688b974f0cd67fd37641d3ce2adb9c696ab09074a966dbbbd0654fd47fc5b57e/merged
|
||||||
@@ -0,0 +1,13 @@
|
|||||||
|
last_run '2026-08-31T02:17:57Z'
|
||||||
|
last_success '2026-08-31T02:17:57Z'
|
||||||
|
last_status 'ok'
|
||||||
|
last_duration '2m52s'
|
||||||
|
schedule 'daily'
|
||||||
|
repo_size_bytes 141959062
|
||||||
|
repo_size_human '135.4 MB'
|
||||||
|
snapshot_count 67
|
||||||
|
stats_known True
|
||||||
|
quota_gb 50
|
||||||
|
last_integrity_check '2026-08-31T08:31:18Z'
|
||||||
|
last_integrity_ok True
|
||||||
|
last_integrity_depth 'structure'
|
||||||
@@ -0,0 +1,8 @@
|
|||||||
|
--- locks BEFORE (observer uses --no-lock) ---
|
||||||
|
lock-count=0
|
||||||
|
--- snapshots for docmost ---
|
||||||
|
ID Time Host Tags Paths
|
||||||
|
--------------------------------------------------------------------------------------------------------------------
|
||||||
|
a07c36a1 2026-08-31 02:17:11 demo-hp felhom-offbox,docmost /mnt/sys_drive/felhom-data/backups/primary/docmost
|
||||||
|
--------------------------------------------------------------------------------------------------------------------
|
||||||
|
1 snapshots
|
||||||
+13
@@ -0,0 +1,13 @@
|
|||||||
|
=== repo totals ===
|
||||||
|
{"total_size":5958072637,"total_file_count":1380,"snapshots_count":67}
|
||||||
|
|
||||||
|
=== per-app latest snapshot: id, size ===
|
||||||
|
docmost a07c36a1 {"total_size":118207270,"total_file_count":20,"snapshots_count":1}
|
||||||
|
romm df7ad253 {"total_size":184706816,"total_file_count":18,"snapshots_count":1}
|
||||||
|
privatebin bb81b9da {"total_size":2123006,"total_file_count":13,"snapshots_count":1}
|
||||||
|
paperless NO-SNAPSHOT
|
||||||
|
opengist b5aa8f9b {"total_size":185664,"total_file_count":13,"snapshots_count":1}
|
||||||
|
kimai e7dcf8fd {"total_size":213231242,"total_file_count":16,"snapshots_count":1}
|
||||||
|
calibre-web ad2ae49c {"total_size":5808703,"total_file_count":38,"snapshots_count":1}
|
||||||
|
bookstack 91154be7 {"total_size":166641768,"total_file_count":19,"snapshots_count":1}
|
||||||
|
filebrowser NO-SNAPSHOT
|
||||||
@@ -0,0 +1,6 @@
|
|||||||
|
observer running; first samples:
|
||||||
|
13:39:17 locks=0
|
||||||
|
ERR
|
||||||
|
=== scratch BEFORE ===
|
||||||
|
ls: cannot access '/mnt/felhom-drives/*/felhom-data/backups/offsite-restore/': No such file or directory
|
||||||
|
949346897920
|
||||||
@@ -0,0 +1,2 @@
|
|||||||
|
t0=2026-08-31T13:40:10.095555740Z
|
||||||
|
post_http=302 post_time=0.010129
|
||||||
@@ -0,0 +1,6 @@
|
|||||||
|
DONE after 2 polls
|
||||||
|
=== controller log, restore window ===
|
||||||
|
2026/08/31 13:40:10 auth.go:134: [DEBUG] [web] auth: valid session for POST /backup/offbox/restore
|
||||||
|
2026/08/31 13:40:10 server.go:393: [DEBUG] [web] ServeHTTP: POST /backup/offbox/restore from 172.18.0.4:34642
|
||||||
|
2026/08/31 13:40:19 offbox_restore.go:298: [INFO] [offbox] restored docmost (a07c36a1, full=false) → /mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost
|
||||||
|
2026/08/31 13:40:19 offbox_handlers.go:372: [INFO] [web] off-box restore docmost completed (full=false, async)
|
||||||
@@ -0,0 +1,33 @@
|
|||||||
|
=== lock observer samples ===
|
||||||
|
13:39:55 rc=0 locks=0 ids=
|
||||||
|
13:39:59 rc=0 locks=0 ids=
|
||||||
|
13:40:03 rc=0 locks=0 ids=
|
||||||
|
13:40:07 rc=0 locks=0 ids=
|
||||||
|
13:40:11 rc=0 locks=0 ids=
|
||||||
|
13:40:15 rc=0 locks=0 ids=
|
||||||
|
13:40:19 rc=0 locks=0 ids=
|
||||||
|
13:40:23 rc=0 locks=0 ids=
|
||||||
|
13:40:27 rc=0 locks=0 ids=
|
||||||
|
|
||||||
|
=== scratch AFTER (unit restore) ===
|
||||||
|
118761552 /mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost
|
||||||
|
/mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost/mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260822T215432Z-docmost-postgres.sql
|
||||||
|
/mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost/mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260822T140924Z-docmost-postgres.sql
|
||||||
|
/mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost/mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260822T140713Z-docmost-postgres.sql
|
||||||
|
/mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost/mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260822T140501Z-docmost-postgres.sql
|
||||||
|
/mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost/mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260822T162708Z-docmost-postgres.sql
|
||||||
|
/mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost/mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260822T135827Z-docmost-postgres.sql
|
||||||
|
/mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost/mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/docmost-postgres.sql
|
||||||
|
/mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost/mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260822T162347Z-docmost-postgres.sql
|
||||||
|
/mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost/mnt/sys_drive/felhom-data/backups/primary/docmost/compose/.felhom.yml
|
||||||
|
/mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost/mnt/sys_drive/felhom-data/backups/primary/docmost/compose/docker-compose.yml
|
||||||
|
/mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost/mnt/sys_drive/felhom-data/backups/primary/docmost/compose/app.yaml
|
||||||
|
/mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost/mnt/sys_drive/felhom-data/backups/primary/docmost/volume-dumps/docmost_docmost_postgres_data.tar
|
||||||
|
/mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost/mnt/sys_drive/felhom-data/backups/primary/docmost/volume-dumps/docmost_docmost_redis_data.tar
|
||||||
|
/mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost/mnt/sys_drive/felhom-data/backups/primary/docmost/volume-dumps/docmost_docmost_storage.tar
|
||||||
|
/mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost/mnt/sys_drive/felhom-data/backups/primary/docmost/manifest.json
|
||||||
|
/mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost/.felhom-restore-complete.json
|
||||||
|
|
||||||
|
file count: 16
|
||||||
|
=== marker ===
|
||||||
|
{"schema":1,"snapshot_id":"a07c36a1","full":false,"finished_at":"2026-08-31T13:40:19Z"}
|
||||||
File diff suppressed because one or more lines are too long
@@ -0,0 +1,5 @@
|
|||||||
|
t0=13:41:28
|
||||||
|
{"data":{"depth":"100%","duration_ms":40257,"ok":true,"read_data_subset":"100%","skip_reason":"","skipped":false,"unreachable":false},"message":"Az ellenőrzés rendben lezajlott","ok":true}
|
||||||
|
|
||||||
|
http=200 time_total=42.701280
|
||||||
|
t1=13:42:11
|
||||||
+1
@@ -0,0 +1 @@
|
|||||||
|
sh: 1: Syntax error: Unterminated quoted string
|
||||||
@@ -0,0 +1,170 @@
|
|||||||
|
13:39:55 rc=0 locks=0 ids=
|
||||||
|
13:39:59 rc=0 locks=0 ids=
|
||||||
|
13:40:03 rc=0 locks=0 ids=
|
||||||
|
13:40:07 rc=0 locks=0 ids=
|
||||||
|
13:40:11 rc=0 locks=0 ids=
|
||||||
|
13:40:15 rc=0 locks=0 ids=
|
||||||
|
13:40:19 rc=0 locks=0 ids=
|
||||||
|
13:40:23 rc=0 locks=0 ids=
|
||||||
|
13:40:27 rc=0 locks=0 ids=
|
||||||
|
13:40:31 rc=0 locks=0 ids=
|
||||||
|
13:40:35 rc=0 locks=0 ids=
|
||||||
|
13:40:39 rc=0 locks=0 ids=
|
||||||
|
13:40:43 rc=0 locks=0 ids=
|
||||||
|
13:40:47 rc=0 locks=0 ids=
|
||||||
|
13:40:50 rc=0 locks=0 ids=
|
||||||
|
13:40:54 rc=0 locks=0 ids=
|
||||||
|
13:40:58 rc=0 locks=0 ids=
|
||||||
|
13:41:02 rc=0 locks=0 ids=
|
||||||
|
13:41:06 rc=0 locks=0 ids=
|
||||||
|
13:41:10 rc=0 locks=0 ids=
|
||||||
|
13:41:14 rc=0 locks=0 ids=
|
||||||
|
13:41:18 rc=0 locks=0 ids=
|
||||||
|
13:41:22 rc=0 locks=0 ids=
|
||||||
|
13:41:26 rc=0 locks=0 ids=
|
||||||
|
13:41:30 rc=0 locks=0 ids=
|
||||||
|
13:41:34 rc=0 locks=1 ids=81fd4d4200d848466e18cb7a9d8e0d4300d95707c68417449178c3023531eecf
|
||||||
|
13:41:38 rc=0 locks=1 ids=81fd4d4200d848466e18cb7a9d8e0d4300d95707c68417449178c3023531eecf
|
||||||
|
13:41:42 rc=0 locks=1 ids=81fd4d4200d848466e18cb7a9d8e0d4300d95707c68417449178c3023531eecf
|
||||||
|
13:41:46 rc=0 locks=1 ids=81fd4d4200d848466e18cb7a9d8e0d4300d95707c68417449178c3023531eecf
|
||||||
|
13:41:50 rc=0 locks=1 ids=81fd4d4200d848466e18cb7a9d8e0d4300d95707c68417449178c3023531eecf
|
||||||
|
13:41:54 rc=0 locks=1 ids=81fd4d4200d848466e18cb7a9d8e0d4300d95707c68417449178c3023531eecf
|
||||||
|
13:41:58 rc=0 locks=1 ids=81fd4d4200d848466e18cb7a9d8e0d4300d95707c68417449178c3023531eecf
|
||||||
|
13:42:02 rc=0 locks=1 ids=81fd4d4200d848466e18cb7a9d8e0d4300d95707c68417449178c3023531eecf
|
||||||
|
13:42:06 rc=0 locks=1 ids=81fd4d4200d848466e18cb7a9d8e0d4300d95707c68417449178c3023531eecf
|
||||||
|
13:42:10 rc=0 locks=1 ids=81fd4d4200d848466e18cb7a9d8e0d4300d95707c68417449178c3023531eecf
|
||||||
|
13:42:14 rc=0 locks=0 ids=
|
||||||
|
13:42:18 rc=0 locks=0 ids=
|
||||||
|
13:42:22 rc=0 locks=0 ids=
|
||||||
|
13:42:26 rc=0 locks=0 ids=
|
||||||
|
13:42:30 rc=0 locks=0 ids=
|
||||||
|
13:42:34 rc=0 locks=0 ids=
|
||||||
|
13:42:38 rc=0 locks=0 ids=
|
||||||
|
13:42:42 rc=0 locks=0 ids=
|
||||||
|
13:42:46 rc=0 locks=0 ids=
|
||||||
|
13:42:49 rc=0 locks=0 ids=
|
||||||
|
13:42:53 rc=0 locks=0 ids=
|
||||||
|
13:42:57 rc=0 locks=0 ids=
|
||||||
|
13:43:01 rc=0 locks=0 ids=
|
||||||
|
13:43:05 rc=0 locks=0 ids=
|
||||||
|
13:43:09 rc=0 locks=0 ids=
|
||||||
|
13:43:13 rc=0 locks=0 ids=
|
||||||
|
13:43:17 rc=0 locks=0 ids=
|
||||||
|
13:43:21 rc=0 locks=0 ids=
|
||||||
|
13:43:25 rc=0 locks=0 ids=
|
||||||
|
13:43:29 rc=0 locks=0 ids=
|
||||||
|
13:43:33 rc=0 locks=0 ids=
|
||||||
|
13:43:37 rc=0 locks=0 ids=
|
||||||
|
13:43:41 rc=0 locks=0 ids=
|
||||||
|
13:43:45 rc=0 locks=0 ids=
|
||||||
|
13:43:49 rc=0 locks=0 ids=
|
||||||
|
13:43:53 rc=0 locks=0 ids=
|
||||||
|
13:43:57 rc=0 locks=0 ids=
|
||||||
|
13:44:01 rc=0 locks=0 ids=
|
||||||
|
13:44:05 rc=0 locks=1 ids=9222fe46cd531ebe2c729af8ef3912bd77bea8cbe72b0527ce7702d5db678c4f
|
||||||
|
13:44:08 rc=0 locks=0 ids=
|
||||||
|
13:44:12 rc=0 locks=0 ids=
|
||||||
|
13:44:16 rc=0 locks=0 ids=
|
||||||
|
13:44:20 rc=0 locks=0 ids=
|
||||||
|
13:44:24 rc=0 locks=0 ids=
|
||||||
|
13:44:28 rc=0 locks=0 ids=
|
||||||
|
13:44:32 rc=0 locks=0 ids=
|
||||||
|
13:44:36 rc=0 locks=0 ids=
|
||||||
|
13:44:40 rc=0 locks=0 ids=
|
||||||
|
13:44:44 rc=0 locks=0 ids=
|
||||||
|
13:44:48 rc=0 locks=1 ids=7a045a2731f7cdbb2aebf000018f040ed4f81d3ede59ed55921975cfbd8d4a61
|
||||||
|
13:44:52 rc=0 locks=0 ids=
|
||||||
|
13:44:56 rc=0 locks=0 ids=
|
||||||
|
13:45:00 rc=0 locks=0 ids=
|
||||||
|
13:45:04 rc=0 locks=0 ids=
|
||||||
|
13:45:08 rc=0 locks=0 ids=
|
||||||
|
13:45:12 rc=0 locks=0 ids=
|
||||||
|
13:45:16 rc=0 locks=0 ids=
|
||||||
|
13:45:20 rc=0 locks=0 ids=
|
||||||
|
13:45:24 rc=0 locks=0 ids=
|
||||||
|
13:45:28 rc=0 locks=0 ids=
|
||||||
|
13:45:31 rc=0 locks=0 ids=
|
||||||
|
13:45:35 rc=0 locks=0 ids=
|
||||||
|
13:45:39 rc=0 locks=0 ids=
|
||||||
|
13:45:43 rc=0 locks=0 ids=
|
||||||
|
13:45:47 rc=0 locks=0 ids=
|
||||||
|
13:45:51 rc=0 locks=0 ids=
|
||||||
|
13:45:55 rc=0 locks=0 ids=
|
||||||
|
13:45:59 rc=0 locks=0 ids=
|
||||||
|
13:46:03 rc=0 locks=0 ids=
|
||||||
|
13:46:07 rc=0 locks=0 ids=
|
||||||
|
13:46:11 rc=0 locks=0 ids=
|
||||||
|
13:46:15 rc=0 locks=0 ids=
|
||||||
|
13:46:19 rc=0 locks=0 ids=
|
||||||
|
13:46:23 rc=0 locks=0 ids=
|
||||||
|
13:46:27 rc=0 locks=0 ids=
|
||||||
|
13:46:31 rc=0 locks=0 ids=
|
||||||
|
13:46:35 rc=0 locks=0 ids=
|
||||||
|
13:46:39 rc=0 locks=0 ids=
|
||||||
|
13:46:43 rc=0 locks=0 ids=
|
||||||
|
13:46:47 rc=0 locks=0 ids=
|
||||||
|
13:46:51 rc=0 locks=0 ids=
|
||||||
|
13:46:54 rc=0 locks=0 ids=
|
||||||
|
13:46:58 rc=0 locks=0 ids=
|
||||||
|
13:47:02 rc=0 locks=0 ids=
|
||||||
|
13:47:06 rc=0 locks=0 ids=
|
||||||
|
13:47:10 rc=0 locks=0 ids=
|
||||||
|
13:47:14 rc=0 locks=0 ids=
|
||||||
|
13:47:18 rc=0 locks=0 ids=
|
||||||
|
13:47:22 rc=0 locks=0 ids=
|
||||||
|
13:47:26 rc=0 locks=0 ids=
|
||||||
|
13:47:30 rc=0 locks=0 ids=
|
||||||
|
13:47:34 rc=0 locks=0 ids=
|
||||||
|
13:47:38 rc=0 locks=0 ids=
|
||||||
|
13:47:42 rc=0 locks=0 ids=
|
||||||
|
13:47:46 rc=0 locks=0 ids=
|
||||||
|
13:47:50 rc=0 locks=0 ids=
|
||||||
|
13:47:54 rc=0 locks=0 ids=
|
||||||
|
13:47:58 rc=0 locks=0 ids=
|
||||||
|
13:48:01 rc=0 locks=0 ids=
|
||||||
|
13:48:05 rc=0 locks=0 ids=
|
||||||
|
13:48:09 rc=0 locks=0 ids=
|
||||||
|
13:48:13 rc=0 locks=0 ids=
|
||||||
|
13:48:17 rc=0 locks=0 ids=
|
||||||
|
13:48:21 rc=0 locks=0 ids=
|
||||||
|
13:48:25 rc=0 locks=0 ids=
|
||||||
|
13:48:29 rc=0 locks=0 ids=
|
||||||
|
13:48:33 rc=0 locks=0 ids=
|
||||||
|
13:48:37 rc=0 locks=0 ids=
|
||||||
|
13:48:41 rc=0 locks=0 ids=
|
||||||
|
13:48:45 rc=0 locks=0 ids=
|
||||||
|
13:48:49 rc=0 locks=0 ids=
|
||||||
|
13:48:53 rc=0 locks=0 ids=
|
||||||
|
13:48:57 rc=0 locks=0 ids=
|
||||||
|
13:49:01 rc=0 locks=0 ids=
|
||||||
|
13:49:05 rc=0 locks=0 ids=
|
||||||
|
13:49:09 rc=0 locks=0 ids=
|
||||||
|
13:49:13 rc=0 locks=0 ids=
|
||||||
|
13:49:17 rc=0 locks=0 ids=
|
||||||
|
13:49:21 rc=0 locks=0 ids=
|
||||||
|
13:49:25 rc=0 locks=0 ids=
|
||||||
|
13:49:29 rc=0 locks=0 ids=
|
||||||
|
13:49:33 rc=0 locks=0 ids=
|
||||||
|
13:49:36 rc=0 locks=0 ids=
|
||||||
|
13:49:40 rc=0 locks=0 ids=
|
||||||
|
13:49:44 rc=0 locks=0 ids=
|
||||||
|
13:49:48 rc=0 locks=0 ids=
|
||||||
|
13:49:52 rc=0 locks=0 ids=
|
||||||
|
13:49:56 rc=0 locks=0 ids=
|
||||||
|
13:50:00 rc=0 locks=0 ids=
|
||||||
|
13:50:04 rc=0 locks=0 ids=
|
||||||
|
13:50:08 rc=0 locks=0 ids=
|
||||||
|
13:50:12 rc=0 locks=0 ids=
|
||||||
|
13:50:16 rc=0 locks=0 ids=
|
||||||
|
13:50:20 rc=0 locks=0 ids=
|
||||||
|
13:50:24 rc=0 locks=0 ids=
|
||||||
|
13:50:28 rc=0 locks=0 ids=
|
||||||
|
13:50:32 rc=0 locks=0 ids=
|
||||||
|
13:50:36 rc=0 locks=0 ids=
|
||||||
|
13:50:40 rc=0 locks=0 ids=
|
||||||
|
13:50:43 rc=0 locks=0 ids=
|
||||||
|
13:50:47 rc=0 locks=0 ids=
|
||||||
|
13:50:51 rc=0 locks=0 ids=
|
||||||
|
13:50:55 rc=0 locks=0 ids=
|
||||||
|
13:50:59 rc=0 locks=0 ids=
|
||||||
|
13:51:03 rc=0 locks=0 ids=
|
||||||
+6
@@ -0,0 +1,6 @@
|
|||||||
|
c3f1b55c {"total_size":83473654,"total_file_count":29,"snapshots_count":1}
|
||||||
|
|
||||||
|
ad2ae49c {"total_size":5808703,"total_file_count":38,"snapshots_count":1}
|
||||||
|
|
||||||
|
e7dcf8fd {"total_size":213231242,"total_file_count":16,"snapshots_count":1}
|
||||||
|
|
||||||
+10
@@ -0,0 +1,10 @@
|
|||||||
|
13:42:54
|
||||||
|
t0=2026-08-31T13:42:54.542010807Z
|
||||||
|
post_http=302 post_time=0.009733
|
||||||
|
2026/08/31 13:42:11 offbox_integrity.go:306: [INFO] [offbox] integrity: check PASSED in 40s (structure, index, and 100% of the pack data re-read)
|
||||||
|
2026/08/31 13:42:54 auth.go:134: [DEBUG] [web] auth: valid session for POST /backup/offbox/restore
|
||||||
|
2026/08/31 13:42:54 server.go:393: [DEBUG] [web] ServeHTTP: POST /backup/offbox/restore from 172.18.0.4:34642
|
||||||
|
2026/08/31 13:43:05 offbox_restore.go:298: [INFO] [offbox] restored kimai (e7dcf8fd, full=true) → /mnt/felhom-drives/hdd_1/backups/offsite-restore/kimai
|
||||||
|
2026/08/31 13:43:05 offbox_handlers.go:372: [INFO] [web] off-box restore kimai completed (full=true, async)
|
||||||
|
---
|
||||||
|
213231328 /mnt/felhom-drives/hdd_1/backups/offsite-restore/kimai
|
||||||
@@ -0,0 +1,58 @@
|
|||||||
|
13:44:39 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:40 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:40 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:40 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:41 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> snapshots latest --tag opengist --json
|
||||||
|
13:44:42 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> snapshots latest --tag opengist --json
|
||||||
|
13:44:42 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> snapshots latest --tag opengist --json
|
||||||
|
13:44:42 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> snapshots latest --tag opengist --json
|
||||||
|
13:44:42 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:43 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> snapshots latest --tag opengist --json
|
||||||
|
13:44:43 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:43 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> snapshots latest --tag opengist --json
|
||||||
|
13:44:43 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:43 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> snapshots latest --tag opengist --json
|
||||||
|
13:44:43 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:44 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:44 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> unlock
|
||||||
|
13:44:44 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> unlock
|
||||||
|
13:44:45 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> unlock
|
||||||
|
13:44:45 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> unlock
|
||||||
|
13:44:46 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> unlock
|
||||||
|
13:44:46 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> restore b5aa8f9b --target /mnt/felhom-drives/hdd_1/backups/offsite-restore/opengist --include /mnt/sys_drive/felhom-data/backups/primary/opengist
|
||||||
|
13:44:46 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> restore b5aa8f9b --target /mnt/felhom-drives/hdd_1/backups/offsite-restore/opengist --include /mnt/sys_drive/felhom-data/backups/primary/opengist
|
||||||
|
13:44:46 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> restore b5aa8f9b --target /mnt/felhom-drives/hdd_1/backups/offsite-restore/opengist --include /mnt/sys_drive/felhom-data/backups/primary/opengist
|
||||||
|
13:44:47 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:47 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> restore b5aa8f9b --target /mnt/felhom-drives/hdd_1/backups/offsite-restore/opengist --include /mnt/sys_drive/felhom-data/backups/primary/opengist
|
||||||
|
13:44:47 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:47 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> restore b5aa8f9b --target /mnt/felhom-drives/hdd_1/backups/offsite-restore/opengist --include /mnt/sys_drive/felhom-data/backups/primary/opengist
|
||||||
|
13:44:47 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:48 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> restore b5aa8f9b --target /mnt/felhom-drives/hdd_1/backups/offsite-restore/opengist --include /mnt/sys_drive/felhom-data/backups/primary/opengist
|
||||||
|
13:44:48 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:48 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> restore b5aa8f9b --target /mnt/felhom-drives/hdd_1/backups/offsite-restore/opengist --include /mnt/sys_drive/felhom-data/backups/primary/opengist
|
||||||
|
13:44:48 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:48 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> restore b5aa8f9b --target /mnt/felhom-drives/hdd_1/backups/offsite-restore/opengist --include /mnt/sys_drive/felhom-data/backups/primary/opengist
|
||||||
|
13:44:49 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> restore b5aa8f9b --target /mnt/felhom-drives/hdd_1/backups/offsite-restore/opengist --include /mnt/sys_drive/felhom-data/backups/primary/opengist
|
||||||
|
13:44:51 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:51 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:51 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:52 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:52 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:55 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:55 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:55 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:56 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:56 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:58 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:58 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:59 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:59 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:45:00 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:45:00 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:45:02 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:45:02 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:45:03 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:45:03 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:45:04 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:45:04 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:45:06 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
@@ -0,0 +1,48 @@
|
|||||||
|
t0=2026-08-31T13:44:41.569463100Z
|
||||||
|
post_http=302 post_time=0.010061
|
||||||
|
=== distinct restic argv observed (timestamps stripped) ===
|
||||||
|
restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> restore b5aa8f9b --target /mnt/felhom-drives/hdd_1/backups/offsite-restore/opengist --include /mnt/sys_drive/felhom-data/backups/primary/opengist
|
||||||
|
restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> snapshots latest --tag opengist --json
|
||||||
|
restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> unlock
|
||||||
|
=== first+last timestamp per verb ===
|
||||||
|
13:44:39 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:40 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:40 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:40 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:41 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> snapshots latest --tag opengist --json
|
||||||
|
13:44:42 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> snapshots latest --tag opengist --json
|
||||||
|
13:44:42 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> snapshots latest --tag opengist --json
|
||||||
|
13:44:42 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> snapshots latest --tag opengist --json
|
||||||
|
13:44:42 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:43 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> snapshots latest --tag opengist --json
|
||||||
|
13:44:43 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:43 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> snapshots latest --tag opengist --json
|
||||||
|
13:44:43 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:43 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> snapshots latest --tag opengist --json
|
||||||
|
13:44:43 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:44 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:44 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> unlock
|
||||||
|
13:44:44 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> unlock
|
||||||
|
13:44:45 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> unlock
|
||||||
|
13:44:45 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> unlock
|
||||||
|
13:44:46 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> unlock
|
||||||
|
13:44:46 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> restore b5aa8f9b --target /mnt/felhom-drives/hdd_1/backups/offsite-restore/opengist --include /mnt/sys_drive/felhom-data/backups/primary/opengist
|
||||||
|
13:44:46 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> restore b5aa8f9b --target /mnt/felhom-drives/hdd_1/backups/offsite-restore/opengist --include /mnt/sys_drive/felhom-data/backups/primary/opengist
|
||||||
|
13:44:46 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> restore b5aa8f9b --target /mnt/felhom-drives/hdd_1/backups/offsite-restore/opengist --include /mnt/sys_drive/felhom-data/backups/primary/opengist
|
||||||
|
13:44:47 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:47 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> restore b5aa8f9b --target /mnt/felhom-drives/hdd_1/backups/offsite-restore/opengist --include /mnt/sys_drive/felhom-data/backups/primary/opengist
|
||||||
|
13:44:47 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:47 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> restore b5aa8f9b --target /mnt/felhom-drives/hdd_1/backups/offsite-restore/opengist --include /mnt/sys_drive/felhom-data/backups/primary/opengist
|
||||||
|
13:44:47 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:48 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> restore b5aa8f9b --target /mnt/felhom-drives/hdd_1/backups/offsite-restore/opengist --include /mnt/sys_drive/felhom-data/backups/primary/opengist
|
||||||
|
13:44:48 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:48 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> restore b5aa8f9b --target /mnt/felhom-drives/hdd_1/backups/offsite-restore/opengist --include /mnt/sys_drive/felhom-data/backups/primary/opengist
|
||||||
|
13:44:48 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:48 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> restore b5aa8f9b --target /mnt/felhom-drives/hdd_1/backups/offsite-restore/opengist --include /mnt/sys_drive/felhom-data/backups/primary/opengist
|
||||||
|
13:44:49 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> restore b5aa8f9b --target /mnt/felhom-drives/hdd_1/backups/offsite-restore/opengist --include /mnt/sys_drive/felhom-data/backups/primary/opengist
|
||||||
|
13:44:51 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:51 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:51 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:52 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
|
13:44:52 restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> list locks --no-lock
|
||||||
@@ -0,0 +1,11 @@
|
|||||||
|
=== A: restore WITHOUT --verify (kimai e7dcf8fd, 213 MB) ===
|
||||||
|
sh: 7: bc: not found
|
||||||
|
rc=0 wall=s bytes=213272202
|
||||||
|
restoring <Snapshot e7dcf8fd of [/mnt/sys_drive/felhom-data/backups/primary/kimai] at 2026-08-31 02:16:30.041113131 +0000 UTC by root@demo-hp> to /mnt/felhom-drives/hdd_1/backups/offsite-restore/.spike-probe/a
|
||||||
|
|
||||||
|
=== B: restore WITH --verify, same snapshot ===
|
||||||
|
sh: 12: bc: not found
|
||||||
|
rc=0 wall=s bytes=213272202
|
||||||
|
restoring <Snapshot e7dcf8fd of [/mnt/sys_drive/felhom-data/backups/primary/kimai] at 2026-08-31 02:16:30.041113131 +0000 UTC by root@demo-hp> to /mnt/felhom-drives/hdd_1/backups/offsite-restore/.spike-probe/b
|
||||||
|
verifying files in /mnt/felhom-drives/hdd_1/backups/offsite-restore/.spike-probe/b
|
||||||
|
finished verifying 7 files in /mnt/felhom-drives/hdd_1/backups/offsite-restore/.spike-probe/b (took 132ms)
|
||||||
+18
@@ -0,0 +1,18 @@
|
|||||||
|
A no-verify rc=0 wall_ms=2990 bytes=213272202
|
||||||
|
B with-verify rc=0 wall_ms=4756 bytes=213272202
|
||||||
|
verifying files in /mnt/felhom-drives/hdd_1/backups/offsite-restore/.spike-probe/b
|
||||||
|
finished verifying 7 files in /mnt/felhom-drives/hdd_1/backups/offsite-restore/.spike-probe/b (took 126ms)
|
||||||
|
|
||||||
|
=== RED-PROOF: corrupt ONE restored byte, keep size AND mtime, re-run restore --verify ===
|
||||||
|
victim=kimai_kimai_db_data.tar size=160331776
|
||||||
|
after edit: size=160331776 (must be unchanged) sha=dd962a1097077e39 orig-sha=81ce628f45418d93
|
||||||
|
re-run rc=0
|
||||||
|
restoring <Snapshot e7dcf8fd of [/mnt/sys_drive/felhom-data/backups/primary/kimai] at 2026-08-31 02:16:30.041113131 +0000 UTC by root@demo-hp> to /mnt/felhom-drives/hdd_1/backups/offsite-restore/.spike-probe/b
|
||||||
|
verifying files in /mnt/felhom-drives/hdd_1/backups/offsite-restore/.spike-probe/b
|
||||||
|
finished verifying 7 files in /mnt/felhom-drives/hdd_1/backups/offsite-restore/.spike-probe/b (took 125ms)
|
||||||
|
|
||||||
|
=== NEGATIVE CONTROL: same command on the UNTOUCHED copy /mnt/felhom-drives/hdd_1/backups/offsite-restore/.spike-probe/a ===
|
||||||
|
rc=0
|
||||||
|
restoring <Snapshot e7dcf8fd of [/mnt/sys_drive/felhom-data/backups/primary/kimai] at 2026-08-31 02:16:30.041113131 +0000 UTC by root@demo-hp> to /mnt/felhom-drives/hdd_1/backups/offsite-restore/.spike-probe/a
|
||||||
|
verifying files in /mnt/felhom-drives/hdd_1/backups/offsite-restore/.spike-probe/a
|
||||||
|
finished verifying 7 files in /mnt/felhom-drives/hdd_1/backups/offsite-restore/.spike-probe/a (took 132ms)
|
||||||
+19
@@ -0,0 +1,19 @@
|
|||||||
|
=== CASE 1: SIZE changed (truncate 1 byte) ===
|
||||||
|
size=160331775
|
||||||
|
rc=0
|
||||||
|
verifying files in /mnt/felhom-drives/hdd_1/backups/offsite-restore/.spike-probe/b
|
||||||
|
finished verifying 7 files in /mnt/felhom-drives/hdd_1/backups/offsite-restore/.spike-probe/b (took 133ms)
|
||||||
|
size after=160331776 sha=81ce628f45418d93
|
||||||
|
|
||||||
|
=== CASE 2: content corrupt + mtime BUMPED ===
|
||||||
|
sha before rerun=51ae3c7762768fe4
|
||||||
|
rc=0
|
||||||
|
verifying files in /mnt/felhom-drives/hdd_1/backups/offsite-restore/.spike-probe/b
|
||||||
|
finished verifying 7 files in /mnt/felhom-drives/hdd_1/backups/offsite-restore/.spike-probe/b (took 128ms)
|
||||||
|
sha after rerun =81ce628f45418d93 orig=81ce628f45418d93
|
||||||
|
|
||||||
|
=== CASE 3: content corrupt, size+mtime PRESERVED, restore into a VIRGIN target ===
|
||||||
|
(this is the shape an unattended test would meet: a fresh restore, then a check)
|
||||||
|
virgin restore+verify rc=0 wall_ms=3438
|
||||||
|
verifying files in /mnt/felhom-drives/hdd_1/backups/offsite-restore/.spike-probe/g
|
||||||
|
finished verifying 7 files in /mnt/felhom-drives/hdd_1/backups/offsite-restore/.spike-probe/g (took 131ms)
|
||||||
+46
@@ -0,0 +1,46 @@
|
|||||||
|
=== does the recovery unit manifest carry hashes? ===
|
||||||
|
{
|
||||||
|
"schema_version": 2,
|
||||||
|
"app_name": "kimai",
|
||||||
|
"display_name": "Kimai",
|
||||||
|
"controller_version": "0.227.1",
|
||||||
|
"created_at": "2026-08-31T02:16:24Z",
|
||||||
|
"drive": "/mnt/sys_drive",
|
||||||
|
"namespace_root": "/mnt/sys_drive/felhom-data",
|
||||||
|
"image_pins": [
|
||||||
|
"kimai/kimai2:apache-2.57.0",
|
||||||
|
"mariadb:11.6"
|
||||||
|
],
|
||||||
|
"secret_env_vars": [
|
||||||
|
"DB_PASSWORD",
|
||||||
|
"ADMIN_PASSWORD"
|
||||||
|
],
|
||||||
|
"data_key_env_vars": null,
|
||||||
|
"secret_source": "portable secrets (data keys, DB passwords, internal signing secrets) are IN this unit's compose/app.yaml (0600); internet-reachable admin logins are NOT, and come from the guest's app.yaml or are regenerated on restore",
|
||||||
|
"config_files": [
|
||||||
|
"docker-compose.yml",
|
||||||
|
".felhom.yml",
|
||||||
|
"app.yaml"
|
||||||
|
],
|
||||||
|
"db_dumps": [
|
||||||
|
"kimai-mariadb.sql"
|
||||||
|
],
|
||||||
|
"volume_dumps": [
|
||||||
|
"kimai_kimai_db_data.tar",
|
||||||
|
"kimai_kimai_var.tar"
|
||||||
|
],
|
||||||
|
"checksums": {
|
||||||
|
".felhom.yml": "95beb458a42391af4bf78429decfbc60537d8d53d61bea68c38687bcc7fa0d23",
|
||||||
|
"app.yaml": "9192dec10e5aae76c2787d0e82ad39115d43f7011519f854bc06534f930c3e9c",
|
||||||
|
"docker-compose.yml": "c25029f8bd16e6252dceda0b4e87917a8b9ce0b9ad6bdbfc82fa8d4f630e6310"
|
||||||
|
},
|
||||||
|
"portable_secret_env_vars": [
|
||||||
|
"DB_PASSWORD"
|
||||||
|
],
|
||||||
|
"o
|
||||||
|
|
||||||
|
=== does restic ls --json expose per-file content hashes? ===
|
||||||
|
{"time":"2026-08-31T02:16:30.041113131Z","parent":"84542ec89b96fd05d504ca33e5f391a959c2e42e94353ca53e2e46929791d576","tree":"00bbec71aced8301be6100c4ac5de6142defbbf0d24b44898a5dfff059e9969f","paths":["/mnt/sys_drive/felhom-data/backups/primary/kimai"],"hostname":"demo-hp","username":"root","tags":["felhom-offbox","kimai"],"id":"e7dcf8fde91a54f655175b60da2e45cfa7b907be78bc9ac5ead5601f3eaf077d","sho
|
||||||
|
{"name":"mnt","type":"dir","path":"/mnt","uid":0,"gid":0,"mode":2147484141,"permissions":"drwxr-xr-x","mtime":"2026-08-21T16:00:39.422914481Z","atime":"2026-08-21T16:00:39.422914481Z","ctime":"2026-08-21T16:00:39.422914481Z","struct_type":"node"}
|
||||||
|
{"name":"sys_drive","type":"dir","path":"/mnt/sys_drive","uid":0,"gid":0,"mode":2147484141,"permissions":"drwxr-xr-x","mtime":"2026-08-21T16:01:03.794262171Z","atime":"2026-08-21T16:01:03.794262171Z","ctime":"2026-08-21T16:01:03.794262171Z","struct_type":"node"}
|
||||||
|
{"name":"felhom-data","type":"dir","path":"/mnt/sys_drive/felhom-data","uid":0,"gid":0,"mode":2147484141,"permissions":"drwxr-xr-x","mtime":"2026-08-21T16:31:03.891800035Z","atime":"2026-08-21T16:31:03.891800035Z","ctime":"2026-08-21T16:31:03.891800035Z","struct_type":"node"}
|
||||||
@@ -0,0 +1,18 @@
|
|||||||
|
=== a FILE node from restic ls --json ===
|
||||||
|
{"name":".felhom.yml","type":"file","path":"/mnt/sys_drive/felhom-data/backups/primary/kimai/compose/.felhom.yml","uid":0,"gid":0,"size":2235,"mode":420,"permissions":"-rw-r--r--","mtime":"2026-08-31T02:16:24.917098021Z","atime":"2026-08-31T02:16:24.917098021Z","ctime":"2026-08-31T02:16:24.918098034Z","struct_type":"node"}
|
||||||
|
{"name":"app.yaml","type":"file","path":"/mnt/sys_drive/felhom-data/backups/primary/kimai/compose/app.yaml","uid":0,"gid":0,"size":488,"mode":384,"permissions":"-rw-------","mtime":"2026-08-31T02:16:24.918098034Z","atime":"2026-08-31T02:16:24.918098034Z","ctime":"2026-08-31T02:16:24.918098034Z","struct_type":"node"}
|
||||||
|
|
||||||
|
=== the full checksums block of the kimai manifest ===
|
||||||
|
"checksums": {
|
||||||
|
".felhom.yml": "95beb458a42391af4bf78429decfbc60537d8d53d61bea68c38687bcc7fa0d23",
|
||||||
|
"app.yaml": "9192dec10e5aae76c2787d0e82ad39115d43f7011519f854bc06534f930c3e9c",
|
||||||
|
"docker-compose.yml": "c25029f8bd16e6252dceda0b4e87917a8b9ce0b9ad6bdbfc82fa8d4f630e6310"
|
||||||
|
},
|
||||||
|
=== files actually in the unit ===
|
||||||
|
48217 /mnt/sys_drive/felhom-data/backups/primary/kimai/db-dumps/kimai-mariadb.sql
|
||||||
|
2235 /mnt/sys_drive/felhom-data/backups/primary/kimai/compose/.felhom.yml
|
||||||
|
2195 /mnt/sys_drive/felhom-data/backups/primary/kimai/compose/docker-compose.yml
|
||||||
|
488 /mnt/sys_drive/felhom-data/backups/primary/kimai/compose/app.yaml
|
||||||
|
160331776 /mnt/sys_drive/felhom-data/backups/primary/kimai/volume-dumps/kimai_kimai_db_data.tar
|
||||||
|
52845056 /mnt/sys_drive/felhom-data/backups/primary/kimai/volume-dumps/kimai_kimai_var.tar
|
||||||
|
1275 /mnt/sys_drive/felhom-data/backups/primary/kimai/manifest.json
|
||||||
@@ -0,0 +1,10 @@
|
|||||||
|
=== restic local cache present? ===
|
||||||
|
1.1M /root/.cache/restic
|
||||||
|
|
||||||
|
=== kimai 213 MB: cached vs --no-cache (the WAN-truth number) ===
|
||||||
|
cached rc=0 ms=3198
|
||||||
|
--no-cache rc=0 ms=5423
|
||||||
|
identical trees: YES
|
||||||
|
|
||||||
|
=== peak scratch bytes for the biggest app ===
|
||||||
|
213272202 /mnt/felhom-drives/hdd_1/backups/offsite-restore/.spike-probe/i
|
||||||
@@ -0,0 +1,9 @@
|
|||||||
|
privatebin snap=bb81b9da logical= 2123006 restored= 2159870 ms= 2628 rc=0
|
||||||
|
opengist snap=b5aa8f9b logical= 185664 restored= 222528 ms= 2253 rc=0
|
||||||
|
calibre-web snap=ad2ae49c logical= 5808703 restored= 5878335 ms= 3642 rc=0
|
||||||
|
paperless-ngx snap=c3f1b55c logical= 83473654 restored= 83547382 ms= 3062 rc=0
|
||||||
|
bookstack snap=91154be7 logical= 166641768 restored= 166682728 ms= 2745 rc=0
|
||||||
|
docmost snap=a07c36a1 logical= 118207270 restored= 118248230 ms= 3676 rc=0
|
||||||
|
romm snap=df7ad253 logical= 184706816 restored= 184747776 ms= 3978 rc=0
|
||||||
|
kimai snap=e7dcf8fd logical= 213231242 restored= 213272202 ms= 3202 rc=0
|
||||||
|
ALL EIGHT, back to back: 25s
|
||||||
+41
@@ -0,0 +1,41 @@
|
|||||||
|
=== every acquireRunning caller (non-test) ===
|
||||||
|
internal/backup/offbox_integrity.go:272: if err := m.acquireRunning(); err != nil {
|
||||||
|
internal/backup/offbox_restore.go:549: if err := m.acquireRunning(); err != nil {
|
||||||
|
internal/backup/offbox_reconstitute.go:491: if err := m.acquireRunning(); err != nil {
|
||||||
|
internal/backup/backup.go:473: if err := m.acquireRunning(); err != nil {
|
||||||
|
internal/backup/backup.go:935:func (m *Manager) AcquireRunningForTest() error { return m.acquireRunning() }
|
||||||
|
internal/backup/backup.go:939:func (m *Manager) acquireRunning() error {
|
||||||
|
internal/backup/shares_restore.go:203: if err := m.acquireRunning(); err != nil {
|
||||||
|
internal/backup/offbox.go:864: if err := m.acquireRunning(); err != nil {
|
||||||
|
internal/backup/tier2_restore.go:183:// do NOT add an acquireRunning() here. A second acquire would refuse the restore it is guarding.
|
||||||
|
internal/backup/tier2_restore.go:347: if err := m.acquireRunning(); err != nil {
|
||||||
|
|
||||||
|
=== does RestoreOffboxScratch acquire it? ===
|
||||||
|
NO acquireRunning in RestoreOffboxScratch
|
||||||
|
|
||||||
|
=== restore_wizard.go:170-190 ===
|
||||||
|
// restoreOpInFlight reports whether a restore op is in flight, FOR DISPLAY.
|
||||||
|
//
|
||||||
|
// **Use this, not `Manager.IsRunning()`.** The Manager carries two different booleans and they are
|
||||||
|
// not interchangeable:
|
||||||
|
//
|
||||||
|
// - `m.running` (read by `IsRunning`) is the CONCURRENCY single-flight. It is acquired *inside*
|
||||||
|
// the restore function, on the background goroutine — and `RestoreOffboxScratch` never acquires
|
||||||
|
// it at all. So for the verification restore and the full-restore preparation — the wizard's two
|
||||||
|
// most-used actions, and the long ones, since they stream from restic — `IsRunning()` is false
|
||||||
|
// for the entire operation.
|
||||||
|
// - `m.opRunning` (read by `RestoreStatus`) is the DISPLAY flag, set synchronously by
|
||||||
|
// `BeginRestoreOp` in the handler *before* the goroutine launches and cleared by `EndRestoreOp`.
|
||||||
|
// It covers all four offsite actions with no start-up window.
|
||||||
|
//
|
||||||
|
// v0.154.0 shipped with `IsRunning()` here, which made the execution step unreachable for
|
||||||
|
// `RestoreOffboxScratch`: the page offered all three intents, with live buttons, while a restore was
|
||||||
|
// downloading — and the progress banner (which polls the op status) contradicted it on the same
|
||||||
|
// screen. Caught by the operator on the first live click-through.
|
||||||
|
func restoreOpInFlight(st backup.RestoreOpStatus) bool {
|
||||||
|
return st.Running
|
||||||
|
}
|
||||||
|
|
||||||
|
// restoreOpBlocked reports whether a NEW restore must be refused right now, and returns the
|
||||||
|
// Hungarian refusal to show. It reads BOTH flags, deliberately:
|
||||||
|
//
|
||||||
@@ -0,0 +1,111 @@
|
|||||||
|
=== scheduled off-site jobs (main.go) ===
|
||||||
|
663: sched.Every("status-refresh", 10*time.Second, func(ctx context.Context) error {
|
||||||
|
666: sched.Every("stack-scan", 2*time.Minute, func(ctx context.Context) error {
|
||||||
|
669: sched.Every("health-probes", 10*time.Second, func(ctx context.Context) error {
|
||||||
|
679: sched.Every("system-health", healthInterval, func(ctx context.Context) error {
|
||||||
|
726: sched.Every("deadapp-check", 30*time.Second, func(ctx context.Context) error {
|
||||||
|
742: sched.Every("ring-spill", 30*time.Second, func(ctx context.Context) error {
|
||||||
|
876: sched.Daily("db-dump", dbLeg, func(ctx context.Context) error {
|
||||||
|
914: sched.Every("offsite-credential-retry", 5*time.Minute, func(ctx context.Context) error {
|
||||||
|
937: sched.Every("backup-cache", 5*time.Minute, func(ctx context.Context) error {
|
||||||
|
962: sched.Daily("tier2-backup", tier2Leg, func(ctx context.Context) error {
|
||||||
|
1095: sched.Daily("offbox-backup", offboxLeg, func(ctx context.Context) error {
|
||||||
|
1109: sched.Daily("offsite-abandon-sweep", "05:10", func(ctx context.Context) error {
|
||||||
|
1139: sched.Daily("offsite-integrity", "06:00", func(ctx context.Context) error {
|
||||||
|
1147: sched.Daily("metrics-prune", "04:00", func(ctx context.Context) error {
|
||||||
|
1198: sched.Daily("fill-watch", "03:30", func(ctx context.Context) error { return fillWatcher.Check() })
|
||||||
|
1228: sched.Every("hub-report", pushInterval, func(ctx context.Context) error {
|
||||||
|
1256: sched.Every("selfupdate-check", checkInterval, func(ctx context.Context) error {
|
||||||
|
1266: sched.Daily("selfupdate-auto", cfg.SelfUpdate.AutoUpdateTime, func(ctx context.Context) error {
|
||||||
|
1293: sched.Daily("asset-sync", cfg.Assets.SyncSchedule, func(ctx context.Context) error {
|
||||||
|
1464: sched.Every("geo-verify", 6*time.Hour, func(ctx context.Context) error {
|
||||||
|
sync_schedule: "05:00"
|
||||||
|
backup:
|
||||||
|
db_dump_schedule: "02:30"
|
||||||
|
enabled: true
|
||||||
|
prune_schedule: weekly
|
||||||
|
restic_password_file: /opt/docker/felhom-controller/data/restic-password
|
||||||
|
restic_schedule: "03:00"
|
||||||
|
retention:
|
||||||
|
keep_daily: 7
|
||||||
|
keep_monthly: 6
|
||||||
|
keep_weekly: 4
|
||||||
|
customer:
|
||||||
|
domain: enkisfelhom.hu
|
||||||
|
email: doodoo21@freemail.hu
|
||||||
|
id: demo-hp
|
||||||
|
name: Demo HP
|
||||||
|
telegram_chat_id: ""
|
||||||
|
git:
|
||||||
|
branch: main
|
||||||
|
--
|
||||||
|
health_check_schedule: "06:00"
|
||||||
|
healthchecks_base: https://status.felhom.eu
|
||||||
|
system_health_interval: 5m
|
||||||
|
thresholds:
|
||||||
|
backup_max_age_hours: 36
|
||||||
|
cpu_warn_percent: 90
|
||||||
|
disk_crit_percent: 90
|
||||||
|
disk_warn_percent: 80
|
||||||
|
memory_warn_percent: 85
|
||||||
|
temperature_warn_celsius: 75
|
||||||
|
internal/backupwindow/backupwindow.go:54:func LegTimes(start string) (db, tier2, offbox string) {
|
||||||
|
internal/backupwindow/backupwindow.go-55- m, err := ParseHHMM(start)
|
||||||
|
internal/backupwindow/backupwindow.go-56- if err != nil {
|
||||||
|
internal/backupwindow/backupwindow.go-57- return "", "", ""
|
||||||
|
internal/backupwindow/backupwindow.go-58- }
|
||||||
|
internal/backupwindow/backupwindow.go-59- return FmtHHMM(m), FmtHHMM(m + tier2OffsetMin), FmtHHMM(m + offboxOffsetMin)
|
||||||
|
internal/backupwindow/backupwindow.go-60-}
|
||||||
|
internal/backupwindow/backupwindow.go-61-
|
||||||
|
internal/backupwindow/backupwindow.go-62-// GateWindow returns the whole-guest backup gate bounds [W+2h, W+6h) as HH:MM strings (for the UI
|
||||||
|
internal/backupwindow/backupwindow.go-63-// "kb. <from>–<to> között" line and the gate-denial log). Empty strings on an invalid start.
|
||||||
|
internal/backupwindow/backupwindow.go-64-func GateWindow(start string) (from, to string) {
|
||||||
|
internal/backupwindow/backupwindow.go-65- m, err := ParseHHMM(start)
|
||||||
|
internal/backupwindow/backupwindow.go-66- if err != nil {
|
||||||
|
internal/backupwindow/backupwindow.go-67- return "", ""
|
||||||
|
internal/backupwindow/backupwindow.go-68- }
|
||||||
|
internal/backupwindow/backupwindow.go-69- return FmtHHMM(m + gateStartMin), FmtHHMM(m + gateEndMin)
|
||||||
|
internal/backupwindow/backupwindow.go-70-}
|
||||||
|
internal/backupwindow/backupwindow.go-71-
|
||||||
|
internal/backupwindow/backupwindow.go-72-// EffectiveWindow resolves the active window by precedence: a valid settings value wins over a valid
|
||||||
|
internal/backupwindow/backupwindow.go-73-// controller.yaml value, which wins over DefaultWindow. An empty or corrupted value simply falls
|
||||||
|
internal/backupwindow/backupwindow.go-74-// through — so a bad settings string degrades to the yaml default rather than breaking scheduling.
|
||||||
|
internal/backupwindow/backupwindow.go-75-func EffectiveWindow(settingsVal, yamlVal string) string {
|
||||||
|
internal/backupwindow/backupwindow.go-76- if Valid(settingsVal) == nil {
|
||||||
|
internal/backupwindow/backupwindow.go-77- return settingsVal
|
||||||
|
internal/backupwindow/backupwindow.go-78- }
|
||||||
|
internal/backupwindow/backupwindow.go-79- if Valid(yamlVal) == nil {
|
||||||
|
internal/backupwindow/backupwindow.go-80- return yamlVal
|
||||||
|
internal/backupwindow/backupwindow.go-81- }
|
||||||
|
internal/backupwindow/backupwindow.go-82- return DefaultWindow
|
||||||
|
internal/backupwindow/backupwindow.go-83-}
|
||||||
|
10:// DefaultWindow is the last-resort window when neither settings nor controller.yaml supplies one.
|
||||||
|
12:const DefaultWindow = "02:30"
|
||||||
|
17: tier2OffsetMin = 60 // tier-2 mirror at W+60m
|
||||||
|
18: offboxOffsetMin = 105 // off-box copy at W+105m
|
||||||
|
2026/08/31 12:33:29 scheduler.go:102: [INFO] [scheduler] Registered periodic job: status-refresh (every 10s)
|
||||||
|
2026/08/31 12:33:29 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="status-refresh" interval=10s totalJobs=1
|
||||||
|
2026/08/31 12:33:29 scheduler.go:102: [INFO] [scheduler] Registered periodic job: stack-scan (every 2m0s)
|
||||||
|
2026/08/31 12:33:29 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="stack-scan" interval=2m0s totalJobs=2
|
||||||
|
2026/08/31 12:33:29 scheduler.go:102: [INFO] [scheduler] Registered periodic job: health-probes (every 10s)
|
||||||
|
2026/08/31 12:33:29 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="health-probes" interval=10s totalJobs=3
|
||||||
|
2026/08/31 12:33:29 scheduler.go:102: [INFO] [scheduler] Registered periodic job: system-health (every 5m0s)
|
||||||
|
2026/08/31 12:33:29 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="system-health" interval=5m0s totalJobs=4
|
||||||
|
2026/08/31 12:33:29 scheduler.go:102: [INFO] [scheduler] Registered periodic job: deadapp-check (every 30s)
|
||||||
|
2026/08/31 12:33:29 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="deadapp-check" interval=30s totalJobs=5
|
||||||
|
2026/08/31 12:33:29 scheduler.go:102: [INFO] [scheduler] Registered periodic job: ring-spill (every 30s)
|
||||||
|
2026/08/31 12:33:29 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="ring-spill" interval=30s totalJobs=6
|
||||||
|
2026/08/31 12:33:29 scheduler.go:132: [INFO] [scheduler] Daily job db-dump scheduled for 2026-09-01 02:30 CEST
|
||||||
|
2026/08/31 12:33:29 scheduler.go:67: [DEBUG] [scheduler] daily job registered: name="db-dump" schedule="02:30" nextRun=2026-09-01T02:30:00+02:00 totalJobs=7
|
||||||
|
2026/08/31 12:33:29 scheduler.go:102: [INFO] [scheduler] Registered periodic job: offsite-credential-retry (every 5m0s)
|
||||||
|
2026/08/31 12:33:29 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="offsite-credential-retry" interval=5m0s totalJobs=8
|
||||||
|
2026/08/31 12:33:29 scheduler.go:102: [INFO] [scheduler] Registered periodic job: backup-cache (every 5m0s)
|
||||||
|
2026/08/31 12:33:29 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="backup-cache" interval=5m0s totalJobs=9
|
||||||
|
2026/08/31 12:33:29 scheduler.go:132: [INFO] [scheduler] Daily job tier2-backup scheduled for 2026-09-01 03:30 CEST
|
||||||
|
2026/08/31 12:33:29 scheduler.go:67: [DEBUG] [scheduler] daily job registered: name="tier2-backup" schedule="03:30" nextRun=2026-09-01T03:30:00+02:00 totalJobs=10
|
||||||
|
2026/08/31 12:33:29 scheduler.go:132: [INFO] [scheduler] Daily job offbox-backup scheduled for 2026-09-01 04:15 CEST
|
||||||
|
2026/08/31 12:33:29 scheduler.go:67: [DEBUG] [scheduler] daily job registered: name="offbox-backup" schedule="04:15" nextRun=2026-09-01T04:15:00+02:00 totalJobs=11
|
||||||
|
2026/08/31 12:33:29 scheduler.go:132: [INFO] [scheduler] Daily job offsite-abandon-sweep scheduled for 2026-09-01 05:10 CEST
|
||||||
|
2026/08/31 12:33:29 scheduler.go:67: [DEBUG] [scheduler] daily job registered: name="offsite-abandon-sweep" schedule="05:10" nextRun=2026-09-01T05:10:00+02:00 totalJobs=12
|
||||||
|
2026/08/31 12:33:29 scheduler.go:132: [INFO] [scheduler] Daily job offsite-integrity scheduled for 2026-09-01 06:00 CEST
|
||||||
|
grep: write error: Broken pipe
|
||||||
@@ -0,0 +1,39 @@
|
|||||||
|
=== LAYER 3: container /tmp ===
|
||||||
|
/tmp: 1: Syntax error: Unterminated quoted string
|
||||||
|
remaining spike files in container: []
|
||||||
|
wc: invalid option -- '"'
|
||||||
|
Try 'wc --help' for more information.
|
||||||
|
restic processes still running:
|
||||||
|
|
||||||
|
=== LAYER 3b: the scratch dirs my probes created ===
|
||||||
|
161146787 /mnt/felhom-drives/hdd_1/backups/offsite-restore/bookstack
|
||||||
|
5808705 /mnt/felhom-drives/hdd_1/backups/offsite-restore/calibre-web
|
||||||
|
118761552 /mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost
|
||||||
|
213231328 /mnt/felhom-drives/hdd_1/backups/offsite-restore/kimai
|
||||||
|
185751 /mnt/felhom-drives/hdd_1/backups/offsite-restore/opengist
|
||||||
|
72345368 /mnt/felhom-drives/hdd_1/backups/offsite-restore/paperless-ngx
|
||||||
|
2123093 /mnt/felhom-drives/hdd_1/backups/offsite-restore/privatebin
|
||||||
|
after: [bookstack
|
||||||
|
calibre-web
|
||||||
|
paperless-ngx]
|
||||||
|
|
||||||
|
=== LAYER 2: guest /root ===
|
||||||
|
remaining in guest: []
|
||||||
|
=== LAYER 1: PVE host /root ===
|
||||||
|
ssh: Could not resolve hostname hp: Name or service not known
|
||||||
|
=== LAYER 1: PVE host demo-hp /root ===
|
||||||
|
host-leftovers-rc=1 (1 = none found = clean)
|
||||||
|
|
||||||
|
=== LAYER 3 re-verified: container ===
|
||||||
|
container /tmp listed above
|
||||||
|
|
||||||
|
=== LAYER 2 re-verified: guest ===
|
||||||
|
guest-leftovers-rc=1
|
||||||
|
|
||||||
|
=== the app still healthy after all this ===
|
||||||
|
felhom-controller Up About an hour (healthy)
|
||||||
|
docmost Up About an hour (healthy)
|
||||||
|
docmost-redis Up About an hour (healthy)
|
||||||
|
docmost-postgres Up About an hour (healthy)
|
||||||
|
romm Up 4 hours (healthy)
|
||||||
|
romm-db Up 4 hours (healthy)
|
||||||
+3
@@ -0,0 +1,3 @@
|
|||||||
|
last_integrity_check": "2026-08-31T13:41:28Z"
|
||||||
|
last_integrity_ok": true
|
||||||
|
last_integrity_depth": "100%"
|
||||||
@@ -225,7 +225,7 @@ unconditional promise, which CAMPAIGN-11 §7 step 7 measured the customer-facing
|
|||||||
| **R-121** | **A BOX's installed agent can sit releases behind the vouched one and nothing notices — the R-120 gate does not cover it.** demo-hp ran agent **0.113.0** while the hub vouched **0.116.0**, through the whole R-116/R-117 arc, and no signal existed on any channel | **READY (S) — NEW 2026-07-30** | — | **Fourth instance of the drift family** (R-111 golden's agent 17 releases behind, R-115 built+deployed but never published, R-120 golden a controller behind — and now installed-vs-vouched on a live box). **Confirmed at source that R-120's gate cannot catch it:** `hub/internal/web/configs.go:1165-1169` compares `goldenVer` against `store.NewestReportedControllerVersion()` — it is a **golden-artifact vs fleet-CONTROLLER** check and says nothing about the agent installed on a box. **`MinAgent` does not cover it either:** it is used to HOLD the controller floor for a box whose agent is too old (`hub/internal/api/handler.go:530-538`, `store.go:1857`) — protective, not an alarm — and demo-hp's 0.113.0 **equalled** `min_agent` 0.113.0, so even a floor comparison was satisfied. **The cost, measured:** R-117's whole subject is the R-113 conjunction, which landed in **0.114.0** — so the designated drill host could not exercise the code under investigation at all, and the R-117 spike had to route every predicate result through an out-of-repo probe built from `main` instead of the installed agent (`audits/SPIKE-r117-bind-liveness-2026-07-30.md` §1, §2.3). Discovered because the R-117 task made bringing the box current an explicit prerequisite. **Fix shape (not implemented):** the hub already receives `AgentVersion` on every host report, and already has semver comparison in Go — the missing piece is a checker comparing reported agent vs the **vouched** agent and surfacing it, operator-tier. Note the honest tension: a box legitimately lags between publish and deploy, so this wants a staleness window rather than an instant alarm | CC |
|
| **R-121** | **A BOX's installed agent can sit releases behind the vouched one and nothing notices — the R-120 gate does not cover it.** demo-hp ran agent **0.113.0** while the hub vouched **0.116.0**, through the whole R-116/R-117 arc, and no signal existed on any channel | **READY (S) — NEW 2026-07-30** | — | **Fourth instance of the drift family** (R-111 golden's agent 17 releases behind, R-115 built+deployed but never published, R-120 golden a controller behind — and now installed-vs-vouched on a live box). **Confirmed at source that R-120's gate cannot catch it:** `hub/internal/web/configs.go:1165-1169` compares `goldenVer` against `store.NewestReportedControllerVersion()` — it is a **golden-artifact vs fleet-CONTROLLER** check and says nothing about the agent installed on a box. **`MinAgent` does not cover it either:** it is used to HOLD the controller floor for a box whose agent is too old (`hub/internal/api/handler.go:530-538`, `store.go:1857`) — protective, not an alarm — and demo-hp's 0.113.0 **equalled** `min_agent` 0.113.0, so even a floor comparison was satisfied. **The cost, measured:** R-117's whole subject is the R-113 conjunction, which landed in **0.114.0** — so the designated drill host could not exercise the code under investigation at all, and the R-117 spike had to route every predicate result through an out-of-repo probe built from `main` instead of the installed agent (`audits/SPIKE-r117-bind-liveness-2026-07-30.md` §1, §2.3). Discovered because the R-117 task made bringing the box current an explicit prerequisite. **Fix shape (not implemented):** the hub already receives `AgentVersion` on every host report, and already has semver comparison in Go — the missing piece is a checker comparing reported agent vs the **vouched** agent and surfacing it, operator-tier. Note the honest tension: a box legitimately lags between publish and deploy, so this wants a staleness window rather than an instant alarm | CC |
|
||||||
| **R-118** | **An absent drive's union row advertises the ROOT filesystem's capacity as its own.** In the absent-state payload the registry-union row reports `total_bytes: 49675956224 / used_bytes: 4584579072` — **byte-identical to the `local` row** (`durable_id: path:/var/lib/vz`, i.e. `pve-root`) in the same response. The real drive is **4 GB** | **READY (XS) — NEW 2026-07-30** | — | Cause: `statfsCapacity(d.MountPath)` (`disks.go:335-338`) statfs's `/mnt/cel`, which with the device gone is a **bare directory on the root filesystem**. `observe.go:176-183`'s comment warns about exactly this trap and guards the Observe path ("*an unmounted removable dir-storage's mountpoint reverts to a bare directory on root … catastrophic DR mis-id*"); **the union path has no equivalent guard.** **Not a DR mis-id** — `durable_id` on that row is still the correct `uuid:…`, so re-attach identity is safe. It is a **false capacity** reaching every consumer of `total_bytes`/`used_fraction` (fill monitors, storage cards): a detached 4 GB drive advertises 46 GiB at 9.2 % used. Same class as `role.go:180-181` — an absent drive's fields decaying to the root filesystem's. Evidence: `audits/DIAG-r116-disks-payload-2026-07-30.md` §12 | CC |
|
| **R-118** | **An absent drive's union row advertises the ROOT filesystem's capacity as its own.** In the absent-state payload the registry-union row reports `total_bytes: 49675956224 / used_bytes: 4584579072` — **byte-identical to the `local` row** (`durable_id: path:/var/lib/vz`, i.e. `pve-root`) in the same response. The real drive is **4 GB** | **READY (XS) — NEW 2026-07-30** | — | Cause: `statfsCapacity(d.MountPath)` (`disks.go:335-338`) statfs's `/mnt/cel`, which with the device gone is a **bare directory on the root filesystem**. `observe.go:176-183`'s comment warns about exactly this trap and guards the Observe path ("*an unmounted removable dir-storage's mountpoint reverts to a bare directory on root … catastrophic DR mis-id*"); **the union path has no equivalent guard.** **Not a DR mis-id** — `durable_id` on that row is still the correct `uuid:…`, so re-attach identity is safe. It is a **false capacity** reaching every consumer of `total_bytes`/`used_fraction` (fill monitors, storage cards): a detached 4 GB drive advertises 46 GiB at 9.2 % used. Same class as `role.go:180-181` — an absent drive's fields decaying to the root filesystem's. Evidence: `audits/DIAG-r116-disks-payload-2026-07-30.md` §12 | CC |
|
||||||
| **R-95** | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY** | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` | CC |
|
| **R-95** | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY** | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` | CC |
|
||||||
| **R-87** | The restic tier is never restore-tested | **READY — RE-RANKED UP 2026-08-03 (R-86 closed)** | — | Design a controller-side test (no scratch-guest analogue transfers). **Most of what this row needed now exists.** R-86 built the piece that was missing: a tier is proved **per archive generation**, on its own rhythm, with the proof recorded as *which archive* — which is exactly the shape a weekly-ish restic tier needs, and the reason this row could not simply reuse the whole-guest scheduler before. What remains is genuinely restic-specific and is NOT a scheduling problem: there is no scratch-guest analogue, so the test has to be a controller-side restore of a bounded sample into a throwaway path, with its own definition of "proved". **Two things to carry over rather than re-derive:** the proof must record the SNAPSHOT it proved (not a timestamp), and the hub's staleness window must learn this tier's rhythm the way `restoreProvenWindow` now does — a restic tier on a weekly cadence lands on the same false-alarm line the flat 7 days did. **And R-95 still applies:** that credential can delete, so a restic restore-test must never be able to write to the repo | CC |
|
| **R-87** | The restic tier is never restore-tested | **READY — SPIKED 2026-08-31, RE-SCOPE PROPOSED (a DECISION for Viktor). Was: READY — RE-RANKED UP 2026-08-03 (R-86 closed)** | — | **SPIKE VERDICT — `audits/SPIKE-restic-restore-test-2026-08-31.md`. Do NOT build this row as written.** **Q1/Q2 measured:** restic is **0.14.0** (the four source comments asserting it are correct); `--verify` DOES exist and is **not a content check** — a one-byte corruption of a restored 160 MB tar with size and mtime preserved **passed clean**, and verify took 131 ms on a 213 MB tree, which cannot be hashing. **Q3:** no reference for "correct" exists — `restic ls --json` carries no content hash in 0.14.0, and the unit manifest hashes **4 918 B of 213 231 242 B** (R-409). **Q4 measured on demo-hp:** one app ≈ 2.3–4.0 s; **all 8 apps / 774 MB logical = 25 s**, against **40.3 s** for the weekly 100% check beside it — a restore-test is CHEAPER than the check. Peak scratch = the app's full logical size (213 MB largest). Cost is dominated by per-snapshot round-trip, not data: 185 KB takes 2.25 s and 213 MB takes 3.20 s. **Q5:** 25 s against a 2m52s nightly backup — skip-if-busy stays right; **but `RestoreOffboxScratch` takes NO `acquireRunning` (R-408)**. **Q6 observed with a positively-controlled lock sampler:** `restic restore` takes **no lock at all**; the product writes anyway because `unlockStale` runs `restic unlock` — a DELETE verb — before every restore (`offbox_restore.go:289`); and `restic check` DOES take a lock (R-407). **R-95's constraint IS satisfiable:** `--no-lock` + skipping `unlockStale` writes nothing, and both mechanisms exist in 0.14.0 and are unused. **Q7, the deciding one:** of R-353/R-354/R-356/R-358/R-403, an unattended scratch-restore would have caught **ONE (R-356)**. **PROPOSED RE-SCOPE:** from *restore-test the tier* to *prove the off-site snapshot still CONTAINS a recoverable unit* — one app a night, restored to scratch, checked against its own `manifest.json` through the existing `unitCarriesData` (`r403_hollow.go:40`), scratch deleted, the SNAPSHOT recorded as the proof. That catches the one thing the weekly check structurally cannot: **`check` proves the stored bytes are the stored bytes, never that we stored the RIGHT thing** — a hollow unit backs up, checks and restores cleanly and recovers nothing (R-403, measured in bytes 2026-08-31). **Viktor decides: build the narrow version, or close this row as answered by R-359.** | CC |
|
||||||
| **R-191** | **Every weekly offsite backup UPLOADS successfully and then FAILS the job on a prune the box is deliberately not allowed to do — on both demo boxes.** Measured on demo-felhom 2026-08-04 06:49–06:53: the upload completed (223 s, 629 MiB of 1.874 GiB, **67.2 % reused incrementally**), then `ERROR: prune 'ct/9201': proxmox-backup-client failed: Error: permission check failed - missing Datastore.Modify\|Datastore.Prune on /datastore/felhom-offsite/demo-felhom` → `ERROR: Backup of VM 9201 failed - error pruning backups` → `TASK ERROR: job errors`. The hub raised `whole_guest_backup_failed` | **CLOSED — SHIPPED 2026-08-04** (installer **1.25.0**; both live boxes corrected) | — | **This is R-89's rule not reaching the config.** R-89 moved PBS pruning SERVER-SIDE — *"boxes set `keep_last: 0`, ep0 runs prune jobs; box tokens stay write-only, never widen the grant"*. The token behaves exactly as designed: it refuses. But **both** demo boxes still arm the offsite tier with `keep_last=2 prune_pbs_allowed=true` (`backup_targets: [{target_id: felhom-pbs, cadence_seconds: 604800, keep_last: 2}]`), so every run asks for a prune that must fail. **The data is SAFE and that is why this is not a P1:** the snapshot lands before the prune is attempted; what is wrong is the job's VERDICT and the weekly operator e-mail it produces. **But it is corrosive in the specific way this project keeps finding:** a backup that reports FAILED while succeeding trains the operator to discount `whole_guest_backup_failed`, which is the same alert that would carry a real one — and it is exactly the failure the R-100 corollary warns about, an alarm whose text is true and whose trigger is not the thing you would act on. **Fix is one config line per box** (`keep_last: 0` on the PBS tier) plus whatever writes it on a fresh install; **deliberately NOT applied in this session** — the session was a runbook with an explicit "change nothing, and if a change appears necessary, stop and report" rule, and a retention field on a live backup tier is not a change to slip into an observation run. **Check before fixing:** whether ep0's prune jobs actually cover these two namespaces, or the snapshots simply accumulate once the box stops asking **THE GATE WAS RUN FIRST, AND IT MATTERED.** Before disabling anything, ep0 was read (read-only, Tier 2): prune jobs `prune-demo-felhom` and `prune-demo-hp` exist on datastore `felhom-offsite`, one per namespace, `schedule 03:30`, `keep-last 2`, comment *"R-82 retention keep-last=2, server-side (box tokens are write-only)"* — and they have run **every day since 2026-07-27: 18 tasks, all `status=OK`**. The newest task log reads `retention options: --ns demo-felhom --max-depth 0 --keep-last 2` / `keep ct/9201/2026-07-27…` / `keep ct/9201/2026-07-28…` / `TASK OK`. Retention happens, and it happens there. **A METHODOLOGICAL WARNING WORTH MORE THAN THE FIX.** Three separate queries said the OPPOSITE — *no prune jobs have ever run* — and **all three were broken instruments**: `worker-type` where the field is `worker_type`; the value `prune` where the worker type is `prunejob`; and `journalctl -u proxmox-backup` where the unit is `proxmox-backup-proxy`. A fourth reading (3 snapshots under keep-last 2) was mis-framed by CC and self-corrected — the third snapshot had landed AFTER that day's 03:30 window. Acting on any of them would have disabled the only pruning ATTEMPT while reporting that nothing prunes: a weekly false alarm traded for unbounded growth on the protected endpoint, invisible for months. **The gate is what caught it, and only because it demanded evidence rather than a verdict.** **Shipped:** installer **1.25.0** writes `keep_last: 0` on the offsite tier (the agent's existing guard `allowPBSPrune = !primary && keep_last > 0` already reads that as *never prune from the box* — no agent change), the justifying paragraph is rewritten to say where retention lives and cite R-89, and `hostinstall_gates.py` asserts it (red-proved: pinning `keep_last: 2` back fails the gate). **Both live boxes corrected in their own config** — `backup tier armed target=felhom-pbs … keep_last=0 … prune_pbs_allowed=false` on demo-felhom and demo-hp, with the local tier untouched at `keep_last=3`. Served over HTTPS at `1.25.0` with `"keep_last":0` in the served bytes. **STILL TO OBSERVE:** the next weekly offsite run completing OK end-to-end. The change removes the failing step; the *schedule* proving it is next week's event, and this row should carry that line when it happens. | CC |
|
| **R-191** | **Every weekly offsite backup UPLOADS successfully and then FAILS the job on a prune the box is deliberately not allowed to do — on both demo boxes.** Measured on demo-felhom 2026-08-04 06:49–06:53: the upload completed (223 s, 629 MiB of 1.874 GiB, **67.2 % reused incrementally**), then `ERROR: prune 'ct/9201': proxmox-backup-client failed: Error: permission check failed - missing Datastore.Modify\|Datastore.Prune on /datastore/felhom-offsite/demo-felhom` → `ERROR: Backup of VM 9201 failed - error pruning backups` → `TASK ERROR: job errors`. The hub raised `whole_guest_backup_failed` | **CLOSED — SHIPPED 2026-08-04** (installer **1.25.0**; both live boxes corrected) | — | **This is R-89's rule not reaching the config.** R-89 moved PBS pruning SERVER-SIDE — *"boxes set `keep_last: 0`, ep0 runs prune jobs; box tokens stay write-only, never widen the grant"*. The token behaves exactly as designed: it refuses. But **both** demo boxes still arm the offsite tier with `keep_last=2 prune_pbs_allowed=true` (`backup_targets: [{target_id: felhom-pbs, cadence_seconds: 604800, keep_last: 2}]`), so every run asks for a prune that must fail. **The data is SAFE and that is why this is not a P1:** the snapshot lands before the prune is attempted; what is wrong is the job's VERDICT and the weekly operator e-mail it produces. **But it is corrosive in the specific way this project keeps finding:** a backup that reports FAILED while succeeding trains the operator to discount `whole_guest_backup_failed`, which is the same alert that would carry a real one — and it is exactly the failure the R-100 corollary warns about, an alarm whose text is true and whose trigger is not the thing you would act on. **Fix is one config line per box** (`keep_last: 0` on the PBS tier) plus whatever writes it on a fresh install; **deliberately NOT applied in this session** — the session was a runbook with an explicit "change nothing, and if a change appears necessary, stop and report" rule, and a retention field on a live backup tier is not a change to slip into an observation run. **Check before fixing:** whether ep0's prune jobs actually cover these two namespaces, or the snapshots simply accumulate once the box stops asking **THE GATE WAS RUN FIRST, AND IT MATTERED.** Before disabling anything, ep0 was read (read-only, Tier 2): prune jobs `prune-demo-felhom` and `prune-demo-hp` exist on datastore `felhom-offsite`, one per namespace, `schedule 03:30`, `keep-last 2`, comment *"R-82 retention keep-last=2, server-side (box tokens are write-only)"* — and they have run **every day since 2026-07-27: 18 tasks, all `status=OK`**. The newest task log reads `retention options: --ns demo-felhom --max-depth 0 --keep-last 2` / `keep ct/9201/2026-07-27…` / `keep ct/9201/2026-07-28…` / `TASK OK`. Retention happens, and it happens there. **A METHODOLOGICAL WARNING WORTH MORE THAN THE FIX.** Three separate queries said the OPPOSITE — *no prune jobs have ever run* — and **all three were broken instruments**: `worker-type` where the field is `worker_type`; the value `prune` where the worker type is `prunejob`; and `journalctl -u proxmox-backup` where the unit is `proxmox-backup-proxy`. A fourth reading (3 snapshots under keep-last 2) was mis-framed by CC and self-corrected — the third snapshot had landed AFTER that day's 03:30 window. Acting on any of them would have disabled the only pruning ATTEMPT while reporting that nothing prunes: a weekly false alarm traded for unbounded growth on the protected endpoint, invisible for months. **The gate is what caught it, and only because it demanded evidence rather than a verdict.** **Shipped:** installer **1.25.0** writes `keep_last: 0` on the offsite tier (the agent's existing guard `allowPBSPrune = !primary && keep_last > 0` already reads that as *never prune from the box* — no agent change), the justifying paragraph is rewritten to say where retention lives and cite R-89, and `hostinstall_gates.py` asserts it (red-proved: pinning `keep_last: 2` back fails the gate). **Both live boxes corrected in their own config** — `backup tier armed target=felhom-pbs … keep_last=0 … prune_pbs_allowed=false` on demo-felhom and demo-hp, with the local tier untouched at `keep_last=3`. Served over HTTPS at `1.25.0` with `"keep_last":0` in the served bytes. **STILL TO OBSERVE:** the next weekly offsite run completing OK end-to-end. The change removes the failing step; the *schedule* proving it is next week's event, and this row should carry that line when it happens. | CC |
|
||||||
| **R-194** | **PVE's permission cache delays every grant-state verdict by an unknown amount, so "the agent can read it" and "the ACL exists" are not the same measurement.** Observed twice while validating R-190's self-repair on demo-felhom 2026-08-04: both ACL rows for `/storage/felhom-backup` were deleted, and `GET /access/permissions` continued to report `Datastore.AllocateSpace` present — for **~40 s** in one run and **~16 minutes** in another. During that window the capability probe reads healthy and the self-repair does not fire | **OPEN** | — | **Why it matters beyond the delay:** it puts a floor under how fast a lost grant can be noticed, it makes any single permission read a lagging indicator, and — the interesting part — **it is a candidate contributor to R-190's own timeline**: a grant removed at an unknown moment could keep working until a cache expiry, which is exactly the shape of *worked at 04:44, refused at 09:24*. That does not explain what removed it, but it may explain when the refusal SURFACED, and the two have been treated as the same instant. **Not a defect in our code** — it is PVE behaviour, and the mitigation already tolerates it (the repair fires on the next probe after the cache clears). **What is worth deciding:** whether the store-grant probe should ALSO consult the storage content listing as a second signal, since that appeared to reflect the loss immediately (`{"data":[]}` while the permission read still said present) — two signals disagreeing is itself information, and today only one of them is read | CC |
|
| **R-194** | **PVE's permission cache delays every grant-state verdict by an unknown amount, so "the agent can read it" and "the ACL exists" are not the same measurement.** Observed twice while validating R-190's self-repair on demo-felhom 2026-08-04: both ACL rows for `/storage/felhom-backup` were deleted, and `GET /access/permissions` continued to report `Datastore.AllocateSpace` present — for **~40 s** in one run and **~16 minutes** in another. During that window the capability probe reads healthy and the self-repair does not fire | **OPEN** | — | **Why it matters beyond the delay:** it puts a floor under how fast a lost grant can be noticed, it makes any single permission read a lagging indicator, and — the interesting part — **it is a candidate contributor to R-190's own timeline**: a grant removed at an unknown moment could keep working until a cache expiry, which is exactly the shape of *worked at 04:44, refused at 09:24*. That does not explain what removed it, but it may explain when the refusal SURFACED, and the two have been treated as the same instant. **Not a defect in our code** — it is PVE behaviour, and the mitigation already tolerates it (the repair fires on the next probe after the cache clears). **What is worth deciding:** whether the store-grant probe should ALSO consult the storage content listing as a second signal, since that appeared to reflect the loss immediately (`{"data":[]}` while the permission read still said present) — two signals disagreeing is itself information, and today only one of them is read | CC |
|
||||||
| **R-200** | **The DR password-injection seam has a handler, a route and tests — and no form.** `POST /backup/offbox/inject-password` is routed (`controller/internal/web/server.go:510`) to `offboxInjectPasswordHandler` (`offbox_handlers.go:174-196`) → `InjectOffboxPassword` (`backup/offbox.go:541`). **No template in the repository contains that path or any form posting to it** (grep over `internal/web/templates/`: one unrelated hit, an XSS comment) | **PLUMBING COMPLETE** (controller **v0.196.0**); **the FORM is not built — still open** | — | **The tenth instance of this project's built-but-never-wired class, and the exact shape `CLAUDE.md` and `felhom.eu/CLAUDE.md`'s seam-wiring rule were written for:** handler tests that POST directly (`offbox_escrow_test.go:167,180`) prove nothing about reachability. To use the only implemented recovery seam today, a person must hand-craft an authenticated POST with a session cookie and CSRF token. **Note the layering while fixing it:** this form takes a **64-hex repo password**, not a recovery code (`offboxRepoPwPattern`, `offbox.go:543`) — they are different secrets at different layers, and the operator's 2026-08-04 ruling asks for a form that takes **R**. Build the R form and treat this one as the operator/DR fallback it was written as, but **ship it with a render test per branch of whatever gate it sits behind**. Source: `audits/RECON-offsite-dr-chain-2026-08-04.md` §3 link 9 **THE DIAGNOSTIC HALF IS DONE AND IT ANSWERED THE QUESTION.** `--recover-offsite-check` is a `docker exec` escape hatch in the shape of `--print-reset-code`: R on **STDIN** (never argv, never `ps`, never shell history, never a transcript), fetch+unseal via the agent, and a verdict of **two sha256 hashes**. **It compares and never installs** — the recovered password is not written to `offbox/repo_password`; a test asserts the data dir is byte-unchanged and its red-proof (adding the install call) fails it. Confirmed live: `repo_password` mtime still `2026-08-03 07:18:02` after the successful check at `2026-08-04 11:49`. **Exit codes are load-bearing** — `0` match, `2` a clean MISMATCH, `1` a step failed; "it failed" and "it worked and disagreed" must never share a status because only one is a finding about the system. A box with no local password reports distinctly (the rebuilt-box shape, where the next step is to INSTALL rather than compare). **WHAT IS NOT BUILT, deliberately:** no card, no form, no preview, no customer-facing text — building an interface on top of a chain nobody had walked is how the preceding three weeks went wrong. **What remains for this row:** link 9 (the recovered password placed so `WriteOffboxSecrets` keeps it) and the customer-facing shape the operator ruled on 2026-08-04 (yell → R form → preview → proceed), which is now priced against a chain that exists rather than one that is assumed **PART 0 SHIPPED 2026-08-04 (v0.196.0):** `--recover-offsite-install` is the sibling of the check — same fetch/unseal path, same STDIN discipline for R — and it **places** the recovered password via `InjectOffboxPassword`. The confirmation is a SECOND invocation (`--confirm-install`): without it, both hashes print and nothing is written, so the operator sees the comparison before a write is possible. Three outcomes named distinctly: **installed** (no local password — the rebuilt-box shape), **unchanged** (identical key, nothing written), **refused** (a DIFFERENT key present — installing would clobber the key the current repository is encrypted under; exit 2, no force offered). It re-reads the file after writing rather than trusting the call. **Red-proof observed:** removing the confirmation gate makes the dry run write the password. The R-persistence test carries a **positive control** (a planted copy found, then removed and not found). **NOT YET EXERCISED AGAINST A LIVE RECOVERY** — the R-201 drill halted before step 9, so this is unit-proven only. **What remains for this row:** the customer-facing shape the operator ruled on 2026-08-04 (yell → recovery-code form → preview → proceed) | CC |
|
| **R-200** | **The DR password-injection seam has a handler, a route and tests — and no form.** `POST /backup/offbox/inject-password` is routed (`controller/internal/web/server.go:510`) to `offboxInjectPasswordHandler` (`offbox_handlers.go:174-196`) → `InjectOffboxPassword` (`backup/offbox.go:541`). **No template in the repository contains that path or any form posting to it** (grep over `internal/web/templates/`: one unrelated hit, an XSS comment) | **PLUMBING COMPLETE** (controller **v0.196.0**); **the FORM is not built — still open** | — | **The tenth instance of this project's built-but-never-wired class, and the exact shape `CLAUDE.md` and `felhom.eu/CLAUDE.md`'s seam-wiring rule were written for:** handler tests that POST directly (`offbox_escrow_test.go:167,180`) prove nothing about reachability. To use the only implemented recovery seam today, a person must hand-craft an authenticated POST with a session cookie and CSRF token. **Note the layering while fixing it:** this form takes a **64-hex repo password**, not a recovery code (`offboxRepoPwPattern`, `offbox.go:543`) — they are different secrets at different layers, and the operator's 2026-08-04 ruling asks for a form that takes **R**. Build the R form and treat this one as the operator/DR fallback it was written as, but **ship it with a render test per branch of whatever gate it sits behind**. Source: `audits/RECON-offsite-dr-chain-2026-08-04.md` §3 link 9 **THE DIAGNOSTIC HALF IS DONE AND IT ANSWERED THE QUESTION.** `--recover-offsite-check` is a `docker exec` escape hatch in the shape of `--print-reset-code`: R on **STDIN** (never argv, never `ps`, never shell history, never a transcript), fetch+unseal via the agent, and a verdict of **two sha256 hashes**. **It compares and never installs** — the recovered password is not written to `offbox/repo_password`; a test asserts the data dir is byte-unchanged and its red-proof (adding the install call) fails it. Confirmed live: `repo_password` mtime still `2026-08-03 07:18:02` after the successful check at `2026-08-04 11:49`. **Exit codes are load-bearing** — `0` match, `2` a clean MISMATCH, `1` a step failed; "it failed" and "it worked and disagreed" must never share a status because only one is a finding about the system. A box with no local password reports distinctly (the rebuilt-box shape, where the next step is to INSTALL rather than compare). **WHAT IS NOT BUILT, deliberately:** no card, no form, no preview, no customer-facing text — building an interface on top of a chain nobody had walked is how the preceding three weeks went wrong. **What remains for this row:** link 9 (the recovered password placed so `WriteOffboxSecrets` keeps it) and the customer-facing shape the operator ruled on 2026-08-04 (yell → R form → preview → proceed), which is now priced against a chain that exists rather than one that is assumed **PART 0 SHIPPED 2026-08-04 (v0.196.0):** `--recover-offsite-install` is the sibling of the check — same fetch/unseal path, same STDIN discipline for R — and it **places** the recovered password via `InjectOffboxPassword`. The confirmation is a SECOND invocation (`--confirm-install`): without it, both hashes print and nothing is written, so the operator sees the comparison before a write is possible. Three outcomes named distinctly: **installed** (no local password — the rebuilt-box shape), **unchanged** (identical key, nothing written), **refused** (a DIFFERENT key present — installing would clobber the key the current repository is encrypted under; exit 2, no force offered). It re-reads the file after writing rather than trusting the call. **Red-proof observed:** removing the confirmation gate makes the dry run write the password. The R-persistence test carries a **positive control** (a planted copy found, then removed and not found). **NOT YET EXERCISED AGAINST A LIVE RECOVERY** — the R-201 drill halted before step 9, so this is unit-proven only. **What remains for this row:** the customer-facing shape the operator ruled on 2026-08-04 (yell → recovery-code form → preview → proceed) | CC |
|
||||||
@@ -455,7 +455,14 @@ there is one ranking to maintain rather than two.
|
|||||||
a drift gate nobody runs and a test that compares a constant to itself. Cheap and worth doing;
|
a drift gate nobody runs and a test that compares a constant to itself. Cheap and worth doing;
|
||||||
not high-consequence, and it blocks nothing.
|
not high-consequence, and it blocks nothing.
|
||||||
3. ~~**R-86**~~ — **CLOSED 2026-08-03**, agent v0.121.0 + hub v0.91.0, proven live on demo-felhom.
|
3. ~~**R-86**~~ — **CLOSED 2026-08-03**, agent v0.121.0 + hub v0.91.0, proven live on demo-felhom.
|
||||||
4. **R-87** — **re-ranked UP**: R-86 built most of what it was waiting for (per-archive due-ness, a
|
4. **R-87** — **SPIKED 2026-08-31 and now a DECISION, not work.** It also spent 2026-08-22..31 in
|
||||||
|
`CLOSED-ITEMS.md` by mistake while this paragraph ranked it fourth and pointed at nothing
|
||||||
|
(R-405). The spike measured it rather than designing it: a scratch restore of all 8 apps costs
|
||||||
|
25 s against the 40.3 s weekly check, but it would have caught ONE of the five drill-found
|
||||||
|
restore defects. **Recommendation: build the NARROW version — prove the snapshot still
|
||||||
|
CONTAINS a recoverable unit — or close the row.** Viktor's call; see
|
||||||
|
`audits/SPIKE-restic-restore-test-2026-08-31.md`. *The 2026-08-03 rationale, kept:*
|
||||||
|
R-86 built most of what it was waiting for (per-archive due-ness, a
|
||||||
proof that names its archive, and a staleness window that learns a tier's rhythm). What is left is
|
proof that names its archive, and a staleness window that learns a tier's rhythm). What is left is
|
||||||
restic-specific — there is no scratch-guest analogue — so it still needs its own design, but it is
|
restic-specific — there is no scratch-guest analogue — so it still needs its own design, but it is
|
||||||
no longer waiting on a scheduling model that did not exist.
|
no longer waiting on a scheduling model that did not exist.
|
||||||
@@ -581,9 +588,12 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
|||||||
| **R-104** | **An interrupted offsite run leaves an exclusive restic lock the existing self-heal cannot reach.** `resticStep` has `unlock --remove-all` (`internal/backup/offbox.go:634-648`) but `ensureOffboxRepo`'s probe fails first, `classifyResticProbe` (`:77-93`) has no lock case → `"other"` → fail-fast; `ClassifyOffsiteFailure` likewise, so the operator is told *„A távoli mentés ismeretlen okból nem sikerült"* for a precisely-known, self-healable condition **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-28, size S, roadmap state `READY — 2026-07-28`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. ****PARTLY STALE, checked against live source 2026-08-22 — migrated as written per the task rule, with the staleness named rather than edited away.** The self-heal this row calls unreachable was built: `resticStep` escalates to `unlock --remove-all` and retries once (`internal/backup/offbox.go:763-768`), and `unlockStale` runs before every off-site run and restore (`:1274`, `offbox_restore.go:261`). Its premise that the probe fails first is also doubtful: the probe is `restic cat config`, a read that takes no lock. **What REMAINS true:** `ClassifyOffsiteFailure` (`offbox.go:179-193`) still has no lock case, so if a lock ever did survive both layers the customer would still be told an unknown reason. Re-rank on that basis, not on the original text.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | **Was C9-F3.** Reachable by any interruption — container restart, OOM, network drop, host reboot mid-backup. The tier stays dead until a human runs `restic unlock --remove-all`. Flips: the offsite row in map §C; `07` §8 row 15 | CC |
|
| **R-104** | **An interrupted offsite run leaves an exclusive restic lock the existing self-heal cannot reach.** `resticStep` has `unlock --remove-all` (`internal/backup/offbox.go:634-648`) but `ensureOffboxRepo`'s probe fails first, `classifyResticProbe` (`:77-93`) has no lock case → `"other"` → fail-fast; `ClassifyOffsiteFailure` likewise, so the operator is told *„A távoli mentés ismeretlen okból nem sikerült"* for a precisely-known, self-healable condition **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-28, size S, roadmap state `READY — 2026-07-28`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. ****PARTLY STALE, checked against live source 2026-08-22 — migrated as written per the task rule, with the staleness named rather than edited away.** The self-heal this row calls unreachable was built: `resticStep` escalates to `unlock --remove-all` and retries once (`internal/backup/offbox.go:763-768`), and `unlockStale` runs before every off-site run and restore (`:1274`, `offbox_restore.go:261`). Its premise that the probe fails first is also doubtful: the probe is `restic cat config`, a read that takes no lock. **What REMAINS true:** `ClassifyOffsiteFailure` (`offbox.go:179-193`) still has no lock case, so if a lock ever did survive both layers the customer would still be told an unknown reason. Re-rank on that basis, not on the original text.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | **Was C9-F3.** Reachable by any interruption — container restart, OOM, network drop, host reboot mid-backup. The tier stays dead until a human runs `restic unlock --remove-all`. Flips: the offsite row in map §C; `07` §8 row 15 | CC |
|
||||||
| **R-105** | **Three hub-held DR records are empty on the entire live fleet.** `hosts.dr_record_json` = `{}` on all 3 hosts; `host_escrow.directive_json` = `{}` on both escrowed hosts; `dr_recipe.host_half.drives` = `[]` on every customer **including two with enrolled data drives** (916 GB USB on demo-felhom, 938 GB NVMe on demo-hp) **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-28, size M, roadmap state `READY — 2026-07-28`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. ****PARTLY FIXED BY ITS OWN UPDATE.** The `drives` third was traced and populated on both demo boxes on 2026-07-28 (the enrolled drives were never PVE storages, so `isUserDataDrive` never saw them). The other two thirds — `hosts.dr_record_json` and `host_escrow.directive_json` — were NOT re-verified this session and are carried as written.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | These are exactly the fields a host-loss recovery reads: `05-hub-architecture.md:175-176,186` names the slim DR record as one of four durable sources; `06-offsite-connectivity.md:148-150` says the escrow upload carried the DR directive; `felhom-agent/internal/dr/plan.go:34-35` makes `PlannedDrive` the re-attach-by-`durable_id` wrong-disk guard. **The three may have different causes** — `isUserDataDrive` (`internal/hub/dr_recipe.go:129-136`) requires type `usb`/`local-dir` **and** a non-empty `DurableID` **and** `MountPath`, and which of the three fails was not traced. Evidence: `architecture/_recovery-inventory-2026-07-28.md` Part D2.3. **UPDATE 2026-07-28 (vzdump-target move): the `drives` third is TRACED and now POPULATED on both demo boxes.** Cause: the enrolled data drives were never PVE storages at all — only agent-generated systemd mounts — so they never entered `report.StorageTargets` and `isUserDataDrive` never saw them. Giving each drive a `dir` storage at its own mountpoint supplied all three required fields at once (type `local-dir`, fs-UUID durable id, mount path), and the recipe now emits `uuid:91d2dc2d-…`/`/mnt/nvme-1tb` on demo-hp and `uuid:47a3361a-…`/`/mnt/hdd_1` o | CC |
|
| **R-105** | **Three hub-held DR records are empty on the entire live fleet.** `hosts.dr_record_json` = `{}` on all 3 hosts; `host_escrow.directive_json` = `{}` on both escrowed hosts; `dr_recipe.host_half.drives` = `[]` on every customer **including two with enrolled data drives** (916 GB USB on demo-felhom, 938 GB NVMe on demo-hp) **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-28, size M, roadmap state `READY — 2026-07-28`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. ****PARTLY FIXED BY ITS OWN UPDATE.** The `drives` third was traced and populated on both demo boxes on 2026-07-28 (the enrolled drives were never PVE storages, so `isUserDataDrive` never saw them). The other two thirds — `hosts.dr_record_json` and `host_escrow.directive_json` — were NOT re-verified this session and are carried as written.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | These are exactly the fields a host-loss recovery reads: `05-hub-architecture.md:175-176,186` names the slim DR record as one of four durable sources; `06-offsite-connectivity.md:148-150` says the escrow upload carried the DR directive; `felhom-agent/internal/dr/plan.go:34-35` makes `PlannedDrive` the re-attach-by-`durable_id` wrong-disk guard. **The three may have different causes** — `isUserDataDrive` (`internal/hub/dr_recipe.go:129-136`) requires type `usb`/`local-dir` **and** a non-empty `DurableID` **and** `MountPath`, and which of the three fails was not traced. Evidence: `architecture/_recovery-inventory-2026-07-28.md` Part D2.3. **UPDATE 2026-07-28 (vzdump-target move): the `drives` third is TRACED and now POPULATED on both demo boxes.** Cause: the enrolled data drives were never PVE storages at all — only agent-generated systemd mounts — so they never entered `report.StorageTargets` and `isUserDataDrive` never saw them. Giving each drive a `dir` storage at its own mountpoint supplied all three required fields at once (type `local-dir`, fs-UUID durable id, mount path), and the recipe now emits `uuid:91d2dc2d-…`/`/mnt/nvme-1tb` on demo-hp and `uuid:47a3361a-…`/`/mnt/hdd_1` o | CC |
|
||||||
| **R-340** | **The new reachability check does not touch the surface that actually failed.** R-339 reports when the hub cannot READ ep0 — but the read it performs is the `usage` op, which is `proxmox-backup-manager` plus `df` over SSH, and therefore rides the **local API daemon**. The 2026-08-18 incident explicitly CLEARED that daemon: `proxmox-backup.service` was healthy throughout, and it was the **HTTPS proxy on 8007** that was wedged with a full accept queue. **So R-339's check would have returned green for all 9 h 37 m of that outage.** It closes the case where ep0 is unreachable *as a host*; it does not close the case that actually happened. This is not a defect in R-339 — it is the honest boundary of what it watches, recorded so a future reader does not mistake a green box gauge for a working off-site tier | **READY (M) — NEW 2026-08-18** | a tenantsync endpoint-script version bump (the op is added on ep0, so it needs the same version-gated rollout `ErrUsageUnsupported` already models) | Add a **health op** to `scripts/felhom-tenantsync.sh` that probes `https://127.0.0.1:8007/` on ep0 and reports the proxy's fd count and listen-queue depth, then surface it as a third signal. **Overlaps the connections spike (R-336's remaining half): both want the same observations from ep0, so whichever runs SECOND must reuse the first's evidence rather than re-measuring a protected machine twice** **REUSE, per this row's own instruction — the connections spike ran FIRST (2026-08-20) and already produced most of what the health op wants; do not re-measure a protected machine a third time.** Available in `audits/evidence-ep0-established-connections-2026-08-20/`: the proxy **fd count** and its type breakdown (`lsof` + `/proc/<pid>/fd`), the **listen-queue depth** (`ss -lnt` — `Recv-Q 0`, `Send-Q 1024`), the **ESTAB/CLOSE-WAIT split**, the **per-peer** connection histogram, a **31-minute persistence diff** of full 4-tuples, and a **46.18 h** slope with Poisson bounds. What the health op would still add beyond these: a loopback `GET https://127.0.0.1:8007/` probe — the observation that distinguished "process problem" from "network problem" on 2026-08-18 and the one thing this spike did NOT take, because it is the surface R-339 cannot see. **And this spike sharpens what the op should report:** a rising **ESTAB** count is the live signal (CLOSE-WAIT was **0**, not merely flat), and per **R-344** the fd ceiling that matters may be the **agent's**, not only ep0's. | CC |
|
| **R-340** | **The new reachability check does not touch the surface that actually failed.** R-339 reports when the hub cannot READ ep0 — but the read it performs is the `usage` op, which is `proxmox-backup-manager` plus `df` over SSH, and therefore rides the **local API daemon**. The 2026-08-18 incident explicitly CLEARED that daemon: `proxmox-backup.service` was healthy throughout, and it was the **HTTPS proxy on 8007** that was wedged with a full accept queue. **So R-339's check would have returned green for all 9 h 37 m of that outage.** It closes the case where ep0 is unreachable *as a host*; it does not close the case that actually happened. This is not a defect in R-339 — it is the honest boundary of what it watches, recorded so a future reader does not mistake a green box gauge for a working off-site tier | **READY (M) — NEW 2026-08-18** | a tenantsync endpoint-script version bump (the op is added on ep0, so it needs the same version-gated rollout `ErrUsageUnsupported` already models) | Add a **health op** to `scripts/felhom-tenantsync.sh` that probes `https://127.0.0.1:8007/` on ep0 and reports the proxy's fd count and listen-queue depth, then surface it as a third signal. **Overlaps the connections spike (R-336's remaining half): both want the same observations from ep0, so whichever runs SECOND must reuse the first's evidence rather than re-measuring a protected machine twice** **REUSE, per this row's own instruction — the connections spike ran FIRST (2026-08-20) and already produced most of what the health op wants; do not re-measure a protected machine a third time.** Available in `audits/evidence-ep0-established-connections-2026-08-20/`: the proxy **fd count** and its type breakdown (`lsof` + `/proc/<pid>/fd`), the **listen-queue depth** (`ss -lnt` — `Recv-Q 0`, `Send-Q 1024`), the **ESTAB/CLOSE-WAIT split**, the **per-peer** connection histogram, a **31-minute persistence diff** of full 4-tuples, and a **46.18 h** slope with Poisson bounds. What the health op would still add beyond these: a loopback `GET https://127.0.0.1:8007/` probe — the observation that distinguished "process problem" from "network problem" on 2026-08-18 and the one thing this spike did NOT take, because it is the surface R-339 cannot see. **And this spike sharpens what the op should report:** a rising **ESTAB** count is the live signal (CLOSE-WAIT was **0**, not merely flat), and per **R-344** the fd ceiling that matters may be the **agent's**, not only ep0's. | CC |
|
||||||
| **R-404** | **DECISION FOR VIKTOR — should a documentation-only push be subject to the golden-currency gate?** `git push --no-verify` has now been used **six times**, each with a recorded reason, because `repo_gates.py --fast` runs `golden_currency_gate.py` on every push to `felhom.eu` including pushes that touch only `documentation/`. **A guard that is correctly bypassed six times is training everyone to bypass it**, and the seventh bypass will be faster to reach for than the sixth | **OPEN — a DECISION, not work. Filed 2026-08-31, deliberately NOT acted on** | — | **FOR narrowing it:** a docs-only push cannot be the push that finishes a release, so scoping the gate to pushes that touch `controller/`, `agent/` or `hub/` code is arguably not a weakening at all — it would fire on exactly the pushes that can create the gap and on no others. It would also end the habit, which is the real cost being paid now. **AGAINST narrowing it:** the gate was earned by a real recurrence — v0.206.0 shipped while the vouched golden carried 0.205.0, and v0.204.0/v0.205.0 before it — and every narrowing of a guard risks the thing it was built for coming back. The gate is also deliberately `--fast` so that BOTH the pre-push hook and CI run it (R-29's census failure); a scoped version must stay in both or it runs in neither. **IF VIKTOR DOES NOTHING:** the bypass stays routine and the count keeps rising; nothing breaks, and the guard quietly stops being one. **Owner: Viktor decides, CC builds. This task did NOT change the gate.** | Viktor |
|
| **R-404** | **DECISION FOR VIKTOR — should a documentation-only push be subject to the golden-currency gate?** `git push --no-verify` has now been used **seven times** (the seventh 2026-08-31, the R-87 spike's records-only Part 1 commit `6e550ae`), each with a recorded reason, because `repo_gates.py --fast` runs `golden_currency_gate.py` on every push to `felhom.eu` including pushes that touch only `documentation/`. **A guard that is correctly bypassed six times is training everyone to bypass it**, and the seventh bypass will be faster to reach for than the sixth | **OPEN — a DECISION, not work. Filed 2026-08-31, deliberately NOT acted on** | — | **FOR narrowing it:** a docs-only push cannot be the push that finishes a release, so scoping the gate to pushes that touch `controller/`, `agent/` or `hub/` code is arguably not a weakening at all — it would fire on exactly the pushes that can create the gap and on no others. It would also end the habit, which is the real cost being paid now. **AGAINST narrowing it:** the gate was earned by a real recurrence — v0.206.0 shipped while the vouched golden carried 0.205.0, and v0.204.0/v0.205.0 before it — and every narrowing of a guard risks the thing it was built for coming back. The gate is also deliberately `--fast` so that BOTH the pre-push hook and CI run it (R-29's census failure); a scoped version must stay in both or it runs in neither. **IF VIKTOR DOES NOTHING:** the bypass stays routine and the count keeps rising; nothing breaks, and the guard quietly stops being one. **Owner: Viktor decides, CC builds. This task did NOT change the gate.** | Viktor |
|
||||||
| **R-405** | **R-87 sat in `CLOSED-ITEMS.md` for nine days while it was still open, and the register's own ranking paragraph ranked it fourth pointing at nothing.** Established from history, not inferred: it was moved by the 2026-08-22 compression sweep, commit `ef6ac6f` (*One register, enforced by a gate; closed work compressed into siblings, R-376..R-378*) — the same commit and the same defect class R-378 records. **R-378 caught six — R-123, R-190, R-214, R-264, R-295, R-352 — and missed a seventh.** R-87 escaped because its state cell read `READY — RE-RANKED UP 2026-08-03 (R-86 closed)`: the leading verdict is `READY` and the word `closed` later in the same cell describes a **different** row. **Count reproduced independently 2026-08-31, and the predicate decides the answer:** matching an open word anywhere in the state cell convicts **three** of 151 rows (R-87, plus R-224 and R-260, both genuinely closed with the words "open"/"OPEN" inside long prose verdicts); matching the **leading verdict** convicts exactly **one**, R-87; matching the whole row convicts **144**. **Fixed this session:** the row is back in `OPEN-ITEMS.md` verbatim from `ef6ac6f^`, next to R-95 where it sat before, and `scripts/closed_register_gate.py` is the 12th gate. Red-proofed both rules and negative-controlled against the pushed pre-fix files, where it convicts R-87 by name. **`R-398` was ALSO in both registers** — a deliberate cross-reference stub — and is now prose beneath the table rather than a row, because a row in both files is what rule 2 convicts on. | **CLOSED 2026-08-31 — corrected + gated in the same session** | R-378 | Nothing further. The gate's four residual holes are named in its docstring; hole 4 is R-406. | CC |
|
| **R-405** | **R-87 sat in `CLOSED-ITEMS.md` for nine days while it was still open, and the register's own ranking paragraph ranked it fourth pointing at nothing.** Established from history, not inferred: it was moved by the 2026-08-22 compression sweep, commit `ef6ac6f` (*One register, enforced by a gate; closed work compressed into siblings, R-376..R-378*) — the same commit and the same defect class R-378 records. **R-378 caught six — R-123, R-190, R-214, R-264, R-295, R-352 — and missed a seventh.** R-87 escaped because its state cell read `READY — RE-RANKED UP 2026-08-03 (R-86 closed)`: the leading verdict is `READY` and the word `closed` later in the same cell describes a **different** row. **Count reproduced independently 2026-08-31, and the predicate decides the answer:** matching an open word anywhere in the state cell convicts **three** of 151 rows (R-87, plus R-224 and R-260, both genuinely closed with the words "open"/"OPEN" inside long prose verdicts); matching the **leading verdict** convicts exactly **one**, R-87; matching the whole row convicts **144**. **Fixed this session:** the row is back in `OPEN-ITEMS.md` verbatim from `ef6ac6f^`, next to R-95 where it sat before, and `scripts/closed_register_gate.py` is the 12th gate. Red-proofed both rules and negative-controlled against the pushed pre-fix files, where it convicts R-87 by name. **`R-398` was ALSO in both registers** — a deliberate cross-reference stub — and is now prose beneath the table rather than a row, because a row in both files is what rule 2 convicts on. | **CLOSED 2026-08-31 — corrected + gated in the same session** | R-378 | Nothing further. The gate's four residual holes are named in its docstring; hole 4 is R-406. | CC |
|
||||||
| **R-406** | **Two unrelated findings in `OPEN-ITEMS.md` share the identifier R-133.** `OPEN-ITEMS.md:267` is *the hub enforces uniqueness on `customer_id` only* (`domain` is `TEXT NOT NULL DEFAULT ''` with no UNIQUE/CHECK); `OPEN-ITEMS.md:273` is *the vaulted break-glass console credential is PLAINTEXT AT REST*. Different subjects, different owners, one number. Found 2026-08-31 while measuring duplicate ids for R-405's gate — **this is the only such collision in either register** (measured: no id appears twice in `CLOSED-ITEMS.md`, and R-88a/R-88b and R-209/R-209a are distinct suffixed ids, not duplicates). **Why the gate does NOT check for it:** a within-register duplicate rule would fail on this pre-existing row, and a registered-but-failing gate refuses every push. **Renumbering is not obviously safe** — `R-133` is cited elsewhere and a blind renumber breaks whichever citation meant the other one. | **OPEN — LOW** | R-405 | Establish which of the two `R-133` citations exist outside the register, then renumber the one with fewer (or none) and add the within-register duplicate rule to `closed_register_gate.py`. Do NOT renumber before grepping the citations. | CC |
|
| **R-406** | **Two unrelated findings in `OPEN-ITEMS.md` share the identifier R-133.** `OPEN-ITEMS.md:267` is *the hub enforces uniqueness on `customer_id` only* (`domain` is `TEXT NOT NULL DEFAULT ''` with no UNIQUE/CHECK); `OPEN-ITEMS.md:273` is *the vaulted break-glass console credential is PLAINTEXT AT REST*. Different subjects, different owners, one number. Found 2026-08-31 while measuring duplicate ids for R-405's gate — **this is the only such collision in either register** (measured: no id appears twice in `CLOSED-ITEMS.md`, and R-88a/R-88b and R-209/R-209a are distinct suffixed ids, not duplicates). **Why the gate does NOT check for it:** a within-register duplicate rule would fail on this pre-existing row, and a registered-but-failing gate refuses every push. **Renumbering is not obviously safe** — `R-133` is cited elsewhere and a blind renumber breaks whichever citation meant the other one. | **OPEN — LOW** | R-405 | Establish which of the two `R-133` citations exist outside the register, then renumber the one with fewer (or none) and add the within-register duplicate rule to `closed_register_gate.py`. Do NOT renumber before grepping the citations. | CC |
|
||||||
|
| **R-407** | **`restic check` DOES write a lock file to the repository, and the comment above it says it never writes.** `offbox_integrity.go:255` reads *"CheckOffboxIntegrity runs one off-site integrity check. It NEVER writes to the repository: `check` is a read verb, and nothing here prunes, forgets, unlocks or backs up."* The three named verbs are correct; the headline is not. **OBSERVED 2026-08-31 on demo-hp**, with a lock sampler that was positively controlled before it was believed: across the product's own integrity run (13:41:28→13:42:11) the repository went `locks=0` → `locks=1 id=81fd4d4200d848466e18cb7a9d8e0d43…` for nine consecutive samples → `locks=0`. The same instrument saw **zero** locks across two restores, so it is not reporting a constant. **Why it matters and why it is LOW rather than ignorable:** R-95's whole constraint is phrased as "must never be able to write to the repo", and a comment stating a guarantee the code does not provide is this project's most-repeated failure — nine instances. The lock itself is correct behaviour and there is no defect in the check; **the defect is the sentence.** | **OPEN — LOW (a comment, not behaviour)** | R-87, R-359 | Correct the sentence in place — say the check takes a repository LOCK and writes nothing else — and pin it with a test, or pass `--no-lock` and make the sentence true. Do NOT delete the sentence: R-360's rule is that a doc comment claiming a guard is why nobody looks for the missing guard. Evidence: `audits/evidence-spike-restic-restore-2026-08-31/16-q6-locks-full.txt`. | CC |
|
||||||
|
| **R-408** | **`RestoreOffboxScratch` takes NO single-writer flag, and the file that depends on that invariant states it as universal.** `offbox_integrity.go:28` reads *"Every off-site operation takes `acquireRunning` for exactly that reason"* — the reason being that `resticStep` escalates to `unlock --remove-all` and is safe only because the in-process mutex proves no sibling operation is live. **Measured 2026-08-31:** `grep -rn 'acquireRunning()'` finds nine non-test callers and `RestoreOffboxScratch` (`offbox_restore.go:234`) is not among them. `restore_wizard.go:174` records the same fact independently — *"`RestoreOffboxScratch` never acquires it at all"* — and the UI works around it with a separate display flag (`opRunning`), so the gap is known at the web layer and unknown at the one that reasons about repository safety. **The exposure today is small and the exposure tomorrow is not:** the web handler's `restoreOpBlocked()` fences the only caller that exists, and restic's restore takes no lock (R-407's sampler), so nothing currently collides. **An unattended restore-test — R-87 — would be the first caller with no web handler in front of it.** | **OPEN — MEDIUM** | R-87, R-359 | Decide ONE way: either `RestoreOffboxScratch` takes `acquireRunning` (and every existing caller is re-checked for the double-acquire refusal `tier2_restore.go:183` warns about), or `offbox_integrity.go:28`'s sentence is corrected to name the exception. Whichever is chosen, **pin it with a test** — this is a comment asserting an invariant with nothing holding it. Do this BEFORE R-87 ships anything. | CC |
|
||||||
|
| **R-409** | **Nothing in the product can vouch for the bytes of a restored recovery unit — the only hash record covers 0.002 % of it.** MEASURED on demo-hp 2026-08-31 against kimai's restored unit: `manifest.json`'s `checksums` object carries sha256 for `.felhom.yml` (2 235 B), `app.yaml` (488 B) and `docker-compose.yml` (2 195 B) — **4 918 bytes of a 213 231 242-byte unit**. The database dump (48 217 B) and the two named-volume tars (160 331 776 B + 52 845 056 B) — the recoverable data, 99.998 % of the bytes — have no recorded hash anywhere. **And nothing else supplies one:** restic 0.14.0's `restore --verify` is a size-and-mtime reconciliation (a size-and-mtime-preserving one-byte corruption of a 160 MB tar passed clean, red-proofed), and `restic ls --json` file nodes in 0.14.0 carry name, size, mode, uid/gid and three timestamps and **no content hash**. **So "the restore produced correct files" is currently unanswerable by any automated means.** **What is NOT claimed here:** `restic check --read-data-subset=100%` already proves the STORE's packs, and the config files that ARE hashed are the ones a wrong-content failure would be hardest to spot in. | **OPEN — MEDIUM** | R-87, R-361 | Cheapest fix, and it is already half-built: extend the capture's `checksums` to cover `db_dumps` and `volume_dumps` — R-361 already computes a canonical dump sha256 to prove itself, so the value exists at capture time. Then a restore-test has a real reference and R-87's narrow version becomes a content check rather than a completeness one. Evidence: `audits/SPIKE-restic-restore-test-2026-08-31.md` §Q2, §Q3. | CC |
|
||||||
|
|
||||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||||
|
|||||||
Reference in New Issue
Block a user