R-101 + F-DIAG closed, F-OPS documented (manual-restore runbook)
This commit is contained in:
+181
@@ -0,0 +1,181 @@
|
||||
# REPORT — R-101 + F-DIAG + F-OPS (2026-07-28)
|
||||
|
||||
Controller **v0.181.0 → v0.182.0**; `felhom.eu` gains the manual-restore runbook (F-OPS) and the
|
||||
OPEN-ITEMS rows. Written as `REPORT-r101.md` so the shared `REPORT.md` is not clobbered.
|
||||
|
||||
## Baselines (reconfirmed, not copied)
|
||||
`felhom-controller 3db8bfb`, `felhom.eu 6b7d516`, `felhom-agent d5c7691` — all = origin/main; the only
|
||||
dirt in `felhom.eu` was a **foreign** `PROMPT-TEMPLATE.md`. Controller **0.181.0** live on both boxes,
|
||||
hub `felhom-hub:0.80.0` ready 1/1.
|
||||
|
||||
---
|
||||
|
||||
## Phase 0
|
||||
|
||||
**The render sites — three dishonest, two already honest.** The spec listed `backups_apps.html:216`
|
||||
as a defect site; it is in fact the one branch that *already* pairs its timestamp with a status badge.
|
||||
The real third site is the `Tier2DestInactive` branch.
|
||||
|
||||
| site | rendered | honest? |
|
||||
|---|---|---|
|
||||
| `:231` **restore confirm dialog** | `Legutóbbi másolat: {{.Tier2LastRun}}` — raw RFC3339, no status | **NO** — the one that matters |
|
||||
| `:195` `Tier2DestDisconnected` | `Utolsó: …`, no status | **NO** |
|
||||
| `:206` `Tier2DestInactive` | `Utolsó: …`, no status | **NO** |
|
||||
| `:217` main configured branch | `Utolsó: …` **+ status badge** | already honest |
|
||||
| `sharing.html:181` | rendered **only** when status=="ok" | already honest |
|
||||
|
||||
`Tier2LastStatus` was already set unconditionally at `handlers.go:1182`, so this was a wording/anchor
|
||||
problem, not a plumbing one. The restore button was gated on `{{if .Tier2LastRun}}`, so **Scenario C
|
||||
was live-reachable**: a tier that had attempted and never succeeded offered a restore and a timestamp.
|
||||
|
||||
**`cd.LastRun` is written on failure** — `recordTier2Failure` (`tier2.go:573-574`) writes it alongside
|
||||
`LastStatus:"error"`. Identical shape to R-100.
|
||||
|
||||
**Legacy state is universal, not an edge case.** All 7 Tier-2 rows across both boxes had `last_run` and
|
||||
no anchor. Scenario E was therefore the *initial state of every customer*, which is what made the
|
||||
legacy marker non-optional.
|
||||
|
||||
---
|
||||
|
||||
## Part 1 — the strings shipped
|
||||
|
||||
| case | string |
|
||||
|---|---|
|
||||
| dialog, normal | `… Legutóbbi sikeres másolat: 2026-07-28 16:43.` |
|
||||
| dialog, newest attempt failed | `… Legutóbbi sikeres másolat: 2026-07-28 16:40. Figyelem: a legutóbbi mentési kísérlet nem sikerült, ezért a visszaállított fájlok ennél régebbiek lehetnek.` |
|
||||
| card | `Utolsó sikeres: 2 perce` |
|
||||
| never succeeded | `Még nincs sikeres másolat` + `Még nincs sikeres másolat, amiből vissza lehetne állítani.` (restore removed) |
|
||||
| **legacy row** | `Utolsó: …` / `Legutóbbi másolat: …` — **today's wording, unchanged**, logged once per stack |
|
||||
|
||||
**Timestamp made human-readable** (agreed): new `fmtTimeStr` renders Budapest-local `2026-07-28 16:40`
|
||||
instead of the raw UTC `2026-07-28T14:40:55Z` a customer was previously asked to reason about.
|
||||
|
||||
**`SuccessTracked` is what makes the legacy case possible at all.** Without it, "row predates the
|
||||
anchor" and "row has an anchor and it is empty" are indistinguishable — both are `LastSuccess==""` —
|
||||
and every existing row would have rendered as never-succeeded on deploy. Legacy rows migrate on first
|
||||
touch: a row whose last known state was `ok` adopts that time (truthful — under the old code that run
|
||||
did succeed); a row whose last state was `error` seeds **nothing**, because the old data evidences no
|
||||
success.
|
||||
|
||||
## Part 2 — the copy-site hazard, and it was in the path
|
||||
|
||||
The three `record*` helpers each built a **whole `CrossDriveBackup` literal**, with a helper re-applying
|
||||
exactly two fields; everything else was zeroed on every status write. Adding `LastSuccess` to that shape
|
||||
would have had `recordTier2Failure` **clear** it — the mirror image of the defect, firing on the *first*
|
||||
failure rather than lying dormant.
|
||||
|
||||
Replaced with **`tier2Update`**, which copies the existing row and overlays the outcome: **compile-safe
|
||||
by construction** — a new field carries over unless deliberately overwritten, so nothing is preserved by
|
||||
a list that can fall out of date. Callers now clear explicitly what a run invalidates, reproducing the
|
||||
old behaviour exactly.
|
||||
|
||||
**Sweep of other rebuild sites:** `SetTier2Preference` mutates in place (safe); `SetCrossDriveConfig(name, nil)`
|
||||
in `api/router.go:774` is a deliberate delete. No others.
|
||||
|
||||
## Part 3 — F-DIAG
|
||||
|
||||
| class | signal it maps to | message head |
|
||||
|---|---|---|
|
||||
| `quota` | the pre-run soft-quota gate | `A távoli mentés nem fért el a tárhelykereten belül` |
|
||||
| `orphaned` | `ErrOffboxOrphaned` sentinel | `A távoli tárhely egy korábbi, már nem elérhető kulccsal készült` |
|
||||
| `no_repo` | restic "unable to open config file" | `A távoli tárhelyen nincs mentési adattár` |
|
||||
| `no_units` | "produced no snapshots" | `Nem volt mit menteni: egyetlen kijelölt alkalmazásnak sem található mentése` |
|
||||
| `transport` | refused/reset/timeout/authn/host-key | `A távoli tárhely nem érhető el (hálózat vagy bejelentkezés)` |
|
||||
| **`unknown`** | anything else | `A távoli mentés ismeretlen okból nem sikerült` |
|
||||
|
||||
The `unknown` class is deliberate: a cause that cannot be told apart where the error is produced is
|
||||
reported as unknown rather than folded into a neighbour.
|
||||
|
||||
**Secrets — and this caught a bug in my own first attempt.** The old message was
|
||||
`"…: " + err.Error()`, carrying the repo reference `sftp:<user>@<host>:<path>` off the box. My first
|
||||
sanitiser regex-matched `sftp:…` and `user@host` and *looked* complete; its own test caught it leaking
|
||||
on `ssh: connect to host <host> port 23: Connection refused` — a bare hostname in neither shape. It now
|
||||
redacts the target's **actual** host/user/repo-path literally, with the regex kept only as a backstop.
|
||||
Guessing at what a secret looks like fails exactly where it matters.
|
||||
|
||||
## Part 4 — F-OPS
|
||||
|
||||
`documentation/runbooks/RUNBOOK-manual-guest-restore.md`. Grounded in the real bind shape read off live
|
||||
guest 9201, not written from memory. Covers: which `mpN` are storage volumes (restored) versus **host
|
||||
binds** (taken as-is on the target); the `mp9` trap — it embeds the **source** VMID, so restoring to a
|
||||
different VMID can bind **another guest's bootstrap credentials**; strip-and-re-add before first boot;
|
||||
the hookscript check; and a positive pre-start verification that asserts every bind path exists rather
|
||||
than accepting "no error". Docs only, by design.
|
||||
|
||||
---
|
||||
|
||||
## Red-proofs — all observed failing
|
||||
|
||||
| # | red-proof | observed failure |
|
||||
|---|---|---|
|
||||
| A | dialog back on the attempt clock | `the dialog does not name the last SUCCESSFUL copy` |
|
||||
| C | gate the restore on `LastRun` again | `a tier that has NEVER succeeded still offers a restore — the dialog would promise a copy that does not exist` |
|
||||
| D | make the caution unconditional | `a HEALTHY tier shows the failed-attempt caution ("nem sikerült")` |
|
||||
| F | clear the anchor on failure | `a FAILED run wiped the success anchor (round 1) — one bad night would read as 'no copy has ever succeeded'` |
|
||||
| + | raw sanitiser | `the repo reference reached the message ("sftp:" leaked)` |
|
||||
|
||||
**F exercises the real `recordTier2Success` → `recordTier2Failure` sequence**, not a modelled copy — the
|
||||
R-100 lesson. The Scenario A/C/D tests **render the production template tree** and assert on the string
|
||||
the customer reads; a test asserting a template variable would prove nothing about wording, which is
|
||||
the defect.
|
||||
|
||||
`go build`, `go vet ./...`, `go test ./...` — 27 packages, `rc=0`; `template_id_gate.py` and
|
||||
`emoji_gate.py` both OK. Run separately from every commit.
|
||||
|
||||
---
|
||||
|
||||
## LIVE on demo-hp — the rendered dialog, which is the deliverable
|
||||
|
||||
**Legacy state** (before any run under v0.182.0) — today's wording, no fright:
|
||||
```
|
||||
Legutóbbi másolat: 2026-07-28 03:30
|
||||
```
|
||||
|
||||
Failure induced genuinely: the Tier-2 destination directory was **moved aside** and replaced by a file,
|
||||
so `mkdir …/recovery-unit` fails. (`chmod` does not work — the controller runs as root, which bypasses
|
||||
permission bits; `chattr +i` is refused, the unprivileged container lacks `CAP_LINUX_IMMUTABLE`. Both
|
||||
were tried and reported rather than glossed.) The real data was only ever moved, never deleted.
|
||||
|
||||
```
|
||||
status = error
|
||||
last_run = 2026-07-28T14:42:18Z ← ADVANCED
|
||||
last_success = 2026-07-28T14:40:55Z ← HELD
|
||||
last_error = mkdir …/paperless-ngx/recovery-unit: …
|
||||
```
|
||||
|
||||
**The rendered dialog, failed state:**
|
||||
```
|
||||
Visszaállítja a hiányzó fájlokat a másodlagos másolatból? A meglévő fájlok NEM módosulnak és NEM
|
||||
törlődnek. Az alkalmazás a művelet idejére leáll. Legutóbbi sikeres másolat: 2026-07-28 16:40.
|
||||
Figyelem: a legutóbbi mentési kísérlet nem sikerült, ezért a visszaállított fájlok ennél régebbiek
|
||||
lehetnek.
|
||||
```
|
||||
|
||||
**The rendered dialog, healthy state** (after restoring the destination and a successful run) — no
|
||||
caution, no tonal change:
|
||||
```
|
||||
Visszaállítja a hiányzó fájlokat a másodlagos másolatból? A meglévő fájlok NEM módosulnak és NEM
|
||||
törlődnek. Az alkalmazás a művelet idejére leáll. Legutóbbi sikeres másolat: 2026-07-28 16:43.
|
||||
```
|
||||
|
||||
Card lines: `Utolsó sikeres: 2 perce` → `Utolsó sikeres: most`.
|
||||
|
||||
**Everything restored:** destination is a directory again, 86 MB intact, mode 755, `.r101-aside` gone,
|
||||
`status=ok`, `last_success=2026-07-28T14:43:23Z`.
|
||||
|
||||
**demo-felhom is the untouched control:** all 5 rows still `tracked=None` after the deploy, rendering
|
||||
today's way, 15/15 containers up. Scenario E holding across a whole box nobody ran.
|
||||
|
||||
---
|
||||
|
||||
## NOT yet live-validated (carried forward)
|
||||
- **F-DIAG's classes** — unit-proven only. No live offsite failure of each class was induced; the
|
||||
`transport` class is the only one this arc exercised indirectly.
|
||||
- **Scenario C live** — the never-succeeded rendering is unit-proven; no fleet row is in that state
|
||||
(every row either migrated or has a real success), and manufacturing one would mean breaking a
|
||||
customer app's only Tier-2 history.
|
||||
- **The Tier-2 restore itself** was not executed — this arc changed what the dialog *says*, not what the
|
||||
restore does.
|
||||
- **R-100's 48h staleness threshold** — injected clock only.
|
||||
- **Fault 4** (restic transport interruption), **R-99**, **F-HUB**, fault 12, the three-way concurrency
|
||||
overlap — next campaign's material, untouched here.
|
||||
@@ -27,7 +27,9 @@ State: `BLOCKED` · `READY` · `WAITING-ON-OPERATOR` · `WATCHING`. Every row ha
|
||||
| **F-CRIT-1** | ~~An app that **fails to restart** after a quiesce never alarms on any channel — `restartAll` discarded the error AND `StateStopped` was whitelisted on invariant I1, which the quiesce path had made false~~ | **SHIPPED + PROVEN-LIVE** (controller v0.179.0, 2026-07-28) | — | Both causes fixed. Live on demo-hp: alarmed 9s after grace expiry, banner shows `(stopped)`; a deliberate user stop stayed silent through 9 dead-app scans | — |
|
||||
| **F-A1** | ~~A restore-test in flight made a healthy backup report as FAILED (HTTP 409 read as a tier failure): breaker armed + operator emailed, on both boxes~~ | **SHIPPED + PROVEN-LIVE** (controller v0.179.0, 2026-07-28) | — | 409 → contention: tier stays DUE, dropped before anything stops (15m), and BLOCKED alarm if contention outlives the agent's 120m ceiling (3h). Hub DB: 409 → **0** operator emails, real failure → **1** | — |
|
||||
| **R-100** | ~~A restic offsite tier that fails every night never goes stale on the hub — `isStale` counted from `LastRun`, which the controller writes unconditionally on failure~~ | **SHIPPED + PROVEN-LIVE** (controller v0.181.0 + hub v0.80.0, 2026-07-28) | — | Anchored on a new `last_success`. **Severity corrected during Phase 0: this was NOT a silence** — `backup_failed` does fire nightly and reaches the operator (live DB: 5 sends). The real defect is **defeated defence in depth**: the hub-side *pull* net was anchored on a field the failing controller keeps refreshing, so it could not compensate for a lost *push* (cf. F-HUB). Live on demo-hp: induced failure → `last_run` advanced 11:25:48Z, `last_success` **held** 11:24:20Z; demo-felhom healthy → anchor advanced. Legacy degrade logged once per customer, live | — |
|
||||
| **R-101** | **Tier-2 (cross-drive) `LastRun` is also written on failure** (`recordTier2Failure`), and three customer surfaces render it **without** a status: the `Tier2DestDisconnected` and `Tier2DestInactive` branches of `backups_apps.html`, and the restore-confirm dialog (`Legutóbbi másolat: {{.Tier2LastRun}}`) — which presents a possibly-failed run's timestamp at the moment the customer decides whether to restore | READY (S) | — | Found by R-100's P0.3/P0.4 sweep; **filed not fixed** (out of that task's path). No hub verdict reads it, so this is a UI-honesty issue, not an alarm one. The two degraded branches do carry a `tag-warn`; the restore-confirm is the sharpest instance. Fix: pair the timestamp with the status everywhere, as the main Tier-2 branch already does | CC |
|
||||
| **R-101** | ~~Tier-2 `LastRun` is written on failure and rendered to the customer as „Legutóbbi másolat" — including in the restore confirm dialog~~ | **SHIPPED + PROVEN-LIVE** (controller v0.182.0, 2026-07-28) | — | `CrossDriveBackup.LastSuccess` + `SuccessTracked`; the dialog names the last **successful** copy and discloses a failed newest attempt. Legacy rows migrate truthfully on first touch (an `ok` row adopts its time; an `error` row seeds nothing) — without the marker all 7 fleet rows would have flipped to „Még nincs sikeres másolat" on deploy. **Part 2:** the three `record*` sites rebuilt the whole struct; replaced by `tier2Update` (copy-and-overlay, safe by construction) — the naive fix would have had `recordTier2Failure` CLEAR the anchor. Live on demo-hp, rendered dialog read in both states | — |
|
||||
| **F-DIAG** | ~~Four distinct offsite failure causes collapse into two operator-visible strings~~ | **SHIPPED** (controller v0.182.0, 2026-07-28) | — | `ClassifyOffsiteFailure` → quota / orphaned / no_repo / no_units / transport / **unknown**, each with its own Hungarian message. Unclassifiable says so rather than being folded into a neighbour. **Secrets:** the old message was a raw `err.Error()` passthrough carrying `sftp:<user>@<host>:<path>`; redaction is now by the target's **actual** host/user/path (a first regex-only attempt leaked on a bare hostname and its own test caught it). Unit-proven; **not** yet exercised by a live offsite failure of each class | — |
|
||||
| **F-OPS** | ~~A manual `pct restore` inherits the source guest's bind mounts — during a real DR, on a different host, under pressure~~ | **DOCUMENTED** (2026-07-28) | — | `documentation/runbooks/RUNBOOK-manual-guest-restore.md`: which `mpN` are volumes vs host binds, the `mp9` source-VMID trap (it can bind **another guest's bootstrap credentials**), strip-and-re-add before first boot, and a positive pre-start verification. Docs only by design — the agent already neutralises binds on its own restore paths, and a second implementation would drift | — |
|
||||
| **F-REBOOT** | ~~A guest rebooted during its backup does not come back — shutdown completes, start never happens, no self-heal; 9m47s total appliance outage with every alarm silent~~ | **SHIPPED + PROVEN-LIVE** (agent v0.107.0, 2026-07-28) | — | 60 s guest-power watchdog; `onboot` is the deliberate-stop discriminator (already the stale-lock path's, and what `pve-guests` consults), retry bounded 3x/1m-2m-4m then escalates once. Live on demo-hp: **120 s unattended** vs the incident's 587 s with a human; Scenario B proven (an `onboot:0` guest left stopped) | — |
|
||||
| **F-LEAK** | ~~A failed restore-test cannot destroy its own scratch guest (403 `VM.Allocate`); the 10-slot VMID band shrinks silently~~ | **SHIPPED + PROVEN-LIVE** (agent v0.110.0 + host-install v1.21.0, 2026-07-28) | — | **Three attempts, two refuted live.** (1) Pool adoption: `PUT /pools/{pool}` also needs `VM.Allocate` on the VM — membership cannot bootstrap its own authority. (2) Per-path `/vms/990000..990009` ACLs: work, but PVE's destroy calls `remove_vm_access` (`LXC.pm:906`) which deletes every ACL at `/vms/<vmid>` — **consumed by the op it authorises**, one use per slot. (3) SHIPPED: 4th root-fenced exception, band enforced in sudoers **literally** (`pct destroy 99000[0-9] --purge`) + in code + at the caller; API destroy still tried first. Live: band PERMITTED, `9201`/`9100`/`9999`/`990010`/`1` REFUSED, and `pct start 990000` REFUSED too | — |
|
||||
| **F-OBS** | ~~`deadapp-check` leaves NO positive observable on a default (info-level) box — "no alarms" was indistinguishable from "never ran"~~ | **SHIPPED + PROVEN-LIVE** (controller v0.180.0 + agent v0.109.0, 2026-07-28) | — | INFO summary every 20th scan carrying scans/evaluated/down. **Agent v0.109.0 fixes the same shape in the guest-power watchdog shipped hours earlier in v0.107.0** — it logged only at startup and when it acted, so its health could be read only from absence | — |
|
||||
|
||||
@@ -0,0 +1,156 @@
|
||||
# RUNBOOK — manual guest restore (`pct restore`), and the bind mounts it drags with it
|
||||
|
||||
**Status:** operational procedure. Written 2026-07-28 to close **F-OPS** (Campaign 8).
|
||||
**Audience:** operator, at a keyboard, usually during a real disaster recovery.
|
||||
|
||||
> **The hazard in one sentence:** `pct restore` recreates the guest from the archive's **own config**,
|
||||
> which includes the source guest's `mpN` bind mounts — so the restored guest can come up bound to a
|
||||
> host path that does not exist on this host, or, worse, to a path that exists and belongs to
|
||||
> **something else**.
|
||||
|
||||
This is not hypothetical and it is not rare: the agent's own restore paths neutralise these mountpoints
|
||||
deliberately (`bindMountOverrides` in `restoretest.go`, `RestoreLXC{MountOverrides:…}` — see
|
||||
`audits/SPIKE-dr-bindmount-source-2026-07-07.md`). A **manual** `pct restore` has no such protection.
|
||||
The automation is safe; the human path is the one that needs this page.
|
||||
|
||||
Timing makes it worse: a manual restore happens during an incident, under pressure, often on a
|
||||
*different* host from the one the archive came from.
|
||||
|
||||
---
|
||||
|
||||
## 1. What a Felhom guest's mounts actually look like
|
||||
|
||||
From live guest 9201 (2026-07-28) — the shape you should expect:
|
||||
|
||||
```
|
||||
rootfs: local-lvm:vm-9201-disk-0,size=32G
|
||||
mp0: local-lvm:vm-9201-disk-1,mp=/var/lib/docker,backup=1,size=50G
|
||||
mp1: local-lvm:vm-9201-disk-2,mp=/mnt/sys_drive,backup=1,size=20G
|
||||
mp8: /mnt/felhom-drives,mp=/mnt/felhom-drives ← HOST BIND
|
||||
mp9: /var/lib/felhom-agent/guests/9201/bootstrap,mp=/etc/felhom-bootstrap,ro=1 ← HOST BIND
|
||||
hookscript: local:snippets/felhom-guest-hook.sh
|
||||
```
|
||||
|
||||
**Two kinds of `mpN`, and only one is dangerous:**
|
||||
|
||||
| kind | looks like | on restore |
|
||||
|---|---|---|
|
||||
| **Volume mount** (`mp0`, `mp1`) | starts with a storage ID — `local-lvm:vm-…` | restored from the archive. Fine. |
|
||||
| **Bind mount** (`mp8`, `mp9`) | starts with an absolute **host path** — `/mnt/…`, `/var/lib/…` | **NOT in the archive.** The path is taken as-is on the target host. |
|
||||
|
||||
The bind mounts are **structural constants** of the Felhom topology, not per-customer data — `mp8` is
|
||||
the shared drive parent, `mp9` is the per-guest bootstrap dir. That is what makes them safe to rewrite
|
||||
by hand: you are restoring a known layout, not guessing.
|
||||
|
||||
**`mp9` is the one that bites.** It embeds the **source** VMID:
|
||||
`/var/lib/felhom-agent/guests/<SOURCE-VMID>/bootstrap`. Restore to a different VMID and the guest
|
||||
either finds nothing there or — if that VMID exists on this host — reads **another guest's bootstrap
|
||||
config, including its credentials**.
|
||||
|
||||
---
|
||||
|
||||
## 2. Before you start — three checks
|
||||
|
||||
Run these on the **target** host, and write the answers down.
|
||||
|
||||
```bash
|
||||
# 1. What does the archive think its mounts are? (read the config WITHOUT restoring)
|
||||
# PBS:
|
||||
proxmox-backup-client restore <snapshot> pct.conf - --repository <repo> 2>/dev/null | grep -E '^(mp|rootfs)'
|
||||
# vzdump tarball:
|
||||
tar -xOf /path/to/vzdump-lxc-<vmid>-<ts>.tar.zst ./etc/vzdump/pct.conf 2>/dev/null | grep -E '^(mp|rootfs)'
|
||||
|
||||
# 2. Does every BIND path exist on THIS host?
|
||||
ls -ld /mnt/felhom-drives /var/lib/felhom-agent/guests/<TARGET-VMID>/bootstrap
|
||||
|
||||
# 3. Is the target VMID free, and is the source VMID something else here?
|
||||
pct list | awk '{print $1}' | grep -x -e <TARGET-VMID> -e <SOURCE-VMID>
|
||||
```
|
||||
|
||||
**Stop and think if:** a bind path is missing, or the source VMID is a live guest on this host.
|
||||
|
||||
---
|
||||
|
||||
## 3. Restore, then fix the binds BEFORE first boot
|
||||
|
||||
Restore **without starting** the guest. `pct restore` does not auto-start, but never pass anything that
|
||||
would, and do not `pct start` until §4 passes.
|
||||
|
||||
```bash
|
||||
pct restore <TARGET-VMID> <volid> --storage <storage> --unprivileged 1
|
||||
```
|
||||
|
||||
Then strip and re-add the binds. Do this even when the VMID is unchanged — it costs seconds and it is
|
||||
the whole point of this page:
|
||||
|
||||
```bash
|
||||
# strip every BIND mountpoint the archive carried (leave the volume mounts mp0/mp1 alone)
|
||||
pct set <TARGET-VMID> --delete mp8
|
||||
pct set <TARGET-VMID> --delete mp9
|
||||
|
||||
# re-add them for THIS host and THIS vmid
|
||||
mkdir -p /var/lib/felhom-agent/guests/<TARGET-VMID>/bootstrap
|
||||
pct set <TARGET-VMID> -mp8 /mnt/felhom-drives,mp=/mnt/felhom-drives
|
||||
pct set <TARGET-VMID> -mp9 /var/lib/felhom-agent/guests/<TARGET-VMID>/bootstrap,mp=/etc/felhom-bootstrap,ro=1
|
||||
```
|
||||
|
||||
> The bootstrap dir must be populated for the guest to enrol. On a host running the agent, the
|
||||
> provisioning back-half owns that directory — prefer letting the agent write it to hand-copying a
|
||||
> `bootstrap.json` between guests. **Never** copy one from another customer's guest: it carries that
|
||||
> customer's tokens.
|
||||
|
||||
Also check the hookscript line: `hookscript: local:snippets/felhom-guest-hook.sh` must point at a
|
||||
snippet that exists **on this host**, or the guest will refuse to start.
|
||||
|
||||
```bash
|
||||
ls -l /var/lib/vz/snippets/felhom-guest-hook.sh || pct set <TARGET-VMID> --delete hookscript
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. Verify before starting — the positive check
|
||||
|
||||
Do not accept "no error" as evidence. Assert what must be **true**:
|
||||
|
||||
```bash
|
||||
# every mpN is either a storage volume or a bind whose host path EXISTS on this host
|
||||
pct config <TARGET-VMID> | grep -E '^mp[0-9]+:' | while read -r line; do
|
||||
src=${line#*: }; src=${src%%,*}
|
||||
case "$src" in
|
||||
/*) [ -e "$src" ] && echo "OK bind $src" || echo "MISSING bind $src <-- fix before start" ;;
|
||||
*) echo "OK volume $src" ;;
|
||||
esac
|
||||
done
|
||||
|
||||
# no mpN may still reference the SOURCE vmid
|
||||
pct config <TARGET-VMID> | grep -E "guests/<SOURCE-VMID>/" && echo "STILL POINTS AT THE SOURCE GUEST"
|
||||
```
|
||||
|
||||
Only when every line reads `OK` and the source-VMID grep is empty:
|
||||
|
||||
```bash
|
||||
pct start <TARGET-VMID>
|
||||
pct exec <TARGET-VMID> -- docker ps --format '{{.Names}} {{.Status}}'
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 5. If the guest is already running with the wrong binds
|
||||
|
||||
Stop it before touching the mounts — a live bind swap under a running container set is how data ends up
|
||||
half-written to two places.
|
||||
|
||||
```bash
|
||||
pct stop <TARGET-VMID>
|
||||
# ...§3 fixes...
|
||||
pct start <TARGET-VMID>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## What this page deliberately does NOT do
|
||||
|
||||
**It does not build tooling.** The agent already neutralises binds on its own restore paths; wrapping
|
||||
the manual path in a script would create a second implementation of that logic, and two implementations
|
||||
of one invariant is how they drift. If the manual path becomes routine, the correct move is to route it
|
||||
through the agent, not to script around it.
|
||||
Reference in New Issue
Block a user