hub v0.129.0: operator raises ONE clean-up window's cap (R-833); restore-beside script + runbooks (R-834)
gates / gates (push) Successful in 30s
gates / gates (push) Successful in 30s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -0,0 +1,14 @@
|
||||
arch: amd64
|
||||
cores: 7
|
||||
features: nesting=1,keyctl=1
|
||||
hookscript: local:snippets/felhom-guest-hook.sh
|
||||
hostname: demo-hp
|
||||
memory: 25898
|
||||
mp0: nvme-scratch:9298/vm-9298-disk-1.raw,mp=/var/lib/felhom,backup=1,size=70G
|
||||
net0: name=eth0,bridge=vmbr0,hwaddr=BC:24:11:0F:7E:C5,ip=dhcp,link_down=1,type=veth
|
||||
net1: name=eth1,bridge=vmbr9,hwaddr=BC:24:11:28:B4:F5,ip=169.254.253.2/30,link_down=1,type=veth
|
||||
onboot: 0
|
||||
ostype: debian
|
||||
rootfs: nvme-scratch:9298/vm-9298-disk-0.raw,size=32G
|
||||
swap: 512
|
||||
unprivileged: 1
|
||||
@@ -0,0 +1,25 @@
|
||||
restore-beside: restoring local:backup/vzdump-lxc-9201-2026_10_04-04_34_55.tar.zst -> 9298 on nvme-scratch with onboot 0 (never started)
|
||||
recovering backed-up configuration from 'local:backup/vzdump-lxc-9201-2026_10_04-04_34_55.tar.zst'
|
||||
Formatting '/mnt/hdd_1/images/9298/vm-9298-disk-0.raw', fmt=raw size=34359738368 preallocation=off
|
||||
Creating filesystem with 8388608 4k blocks and 2097152 inodes
|
||||
Filesystem UUID: 7ddc4545-38c5-4362-8e3d-31a715b29c1b
|
||||
Superblock backups stored on blocks:
|
||||
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
|
||||
4096000, 7962624
|
||||
Formatting '/mnt/hdd_1/images/9298/vm-9298-disk-1.raw', fmt=raw size=75161927680 preallocation=off
|
||||
Creating filesystem with 18350080 4k blocks and 4587520 inodes
|
||||
Filesystem UUID: 28215aa3-d7b4-484b-b2dc-9f85265e80a5
|
||||
Superblock backups stored on blocks:
|
||||
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
|
||||
4096000, 7962624, 11239424
|
||||
restoring 'local:backup/vzdump-lxc-9201-2026_10_04-04_34_55.tar.zst' now..
|
||||
extracting archive '/var/lib/vz/dump/vzdump-lxc-9201-2026_10_04-04_34_55.tar.zst'
|
||||
tar: ./var/lib/felhom/docker/volumes/kimai_kimai_var/_data/cache/prod/pools/app/vl6nzGcpxe/7/D/WpVn7g-RmFPG2mIs6Gaw: time stamp 2026-10-05 04:16:20 is 69996.875442074 s in the future
|
||||
Total bytes read: 15959541760 (15GiB, 276MiB/s)
|
||||
merging backed-up and given configuration..
|
||||
restore-beside: removing host-path binds: mp8,mp9
|
||||
restore-beside: link down: net0
|
||||
restore-beside: link down: net1
|
||||
restore-beside: OK: 9298 has onboot 0, no host-path binds, every NIC link_down; it was not started
|
||||
restore-beside: read it with: pct mount 9298 (then pct unmount 9298; pct destroy 9298 --purge when done)
|
||||
rc=0
|
||||
@@ -0,0 +1,54 @@
|
||||
=== 2026-10-04T06:44:15Z vmid 990000 status=status: stopped
|
||||
lock: create
|
||||
=== 2026-10-04T06:44:19Z vmid 990000 status=status: stopped
|
||||
lock: create
|
||||
=== 2026-10-04T06:44:23Z vmid 990000 status=status: stopped
|
||||
lock: create
|
||||
=== 2026-10-04T06:44:27Z vmid 990000 status=status: stopped
|
||||
lock: create
|
||||
=== 2026-10-04T06:44:31Z vmid 990000 status=status: stopped
|
||||
lock: create
|
||||
=== 2026-10-04T06:44:36Z vmid 990000 status=status: stopped
|
||||
lock: create
|
||||
=== 2026-10-04T06:44:40Z vmid 990000 status=status: stopped
|
||||
lock: create
|
||||
=== 2026-10-04T06:44:44Z vmid 990000 status=status: stopped
|
||||
lock: create
|
||||
=== 2026-10-04T06:44:48Z vmid 990000 status=status: stopped
|
||||
lock: create
|
||||
=== 2026-10-04T06:44:52Z vmid 990000 status=status: stopped
|
||||
lock: create
|
||||
=== 2026-10-04T06:44:56Z vmid 990000 status=status: stopped
|
||||
lock: create
|
||||
=== 2026-10-04T06:45:00Z vmid 990000 status=status: stopped
|
||||
lock: create
|
||||
=== 2026-10-04T06:45:04Z vmid 990000 status=status: stopped
|
||||
lock: create
|
||||
=== 2026-10-04T06:45:08Z vmid 990000 status=status: stopped
|
||||
lock: create
|
||||
=== 2026-10-04T06:45:12Z vmid 990000 status=status: stopped
|
||||
lock: create
|
||||
=== 2026-10-04T06:45:17Z vmid 990000 status=status: stopped
|
||||
hostname: demo-hp
|
||||
mp0: nvme-scratch:990000/vm-990000-disk-1.raw,mp=/var/lib/felhom,backup=1,size=70G
|
||||
mp8: nvme-scratch:990000/vm-990000-disk-2.raw,mp=/mnt/felhom-drives,backup=0,size=1G
|
||||
mp9: nvme-scratch:990000/vm-990000-disk-3.raw,mp=/etc/felhom-bootstrap,backup=0,size=1G
|
||||
net0: name=eth0,bridge=vmbr0,hwaddr=BC:24:11:0F:7E:C5,ip=dhcp,type=veth
|
||||
net1: name=eth1,bridge=vmbr9,hwaddr=BC:24:11:28:B4:F5,ip=169.254.253.2/30,type=veth
|
||||
onboot: 0
|
||||
=== 2026-10-04T06:45:20Z vmid 990000 status=status: running
|
||||
hostname: demo-hp
|
||||
mp0: nvme-scratch:990000/vm-990000-disk-1.raw,mp=/var/lib/felhom,backup=1,size=70G
|
||||
mp8: nvme-scratch:990000/vm-990000-disk-2.raw,mp=/mnt/felhom-drives,backup=0,size=1G
|
||||
mp9: nvme-scratch:990000/vm-990000-disk-3.raw,mp=/etc/felhom-bootstrap,backup=0,size=1G
|
||||
net0: name=eth0,bridge=vmbr0,hwaddr=BC:24:11:0F:7E:C5,ip=dhcp,link_down=1,type=veth
|
||||
net1: name=eth1,bridge=vmbr9,hwaddr=BC:24:11:28:B4:F5,ip=169.254.253.2/30,link_down=1,type=veth
|
||||
onboot: 0
|
||||
=== 2026-10-04T06:45:25Z vmid 990000 status=status: running
|
||||
hostname: demo-hp
|
||||
mp0: nvme-scratch:990000/vm-990000-disk-1.raw,mp=/var/lib/felhom,backup=1,size=70G
|
||||
mp8: nvme-scratch:990000/vm-990000-disk-2.raw,mp=/mnt/felhom-drives,backup=0,size=1G
|
||||
mp9: nvme-scratch:990000/vm-990000-disk-3.raw,mp=/etc/felhom-bootstrap,backup=0,size=1G
|
||||
net0: name=eth0,bridge=vmbr0,hwaddr=BC:24:11:0F:7E:C5,ip=dhcp,link_down=1,type=veth
|
||||
net1: name=eth1,bridge=vmbr9,hwaddr=BC:24:11:28:B4:F5,ip=169.254.253.2/30,link_down=1,type=veth
|
||||
onboot: 0
|
||||
@@ -0,0 +1,27 @@
|
||||
=== felhom-agent 0.138.0 selftest=restore-test ===
|
||||
--- recover: reaping any leaked scratch from a prior crashed test ---
|
||||
recover: examined=0 scratch_destroyed=0 scratch_clean=0
|
||||
restoring local:backup/vzdump-lxc-9201-2026_10_04-04_34_55.tar.zst into scratch band [990000,990009] on nvme-scratch …
|
||||
time=2026-10-04T08:44:13.035+02:00 level=INFO msg="restore-test: space preflight passed" storage=nvme-scratch required_bytes=24520159232 avail_bytes=876487208960
|
||||
time=2026-10-04T08:44:13.137+02:00 level=INFO msg="restore-test: full-fidelity restore params derived from the archive config" scratch=990000 params=4
|
||||
time=2026-10-04T08:45:25.769+02:00 level=INFO msg="audit: gate decision" class=guest_destroy host=demo-hp-bb76ea guest=990000 source=one_shot_job disposition=benign allowed=true reason=benign key_id="" nonce="" durable_id=""
|
||||
time=2026-10-04T08:45:25.769+02:00 level=INFO msg="gate decision" class=guest_destroy guest=990000 source=one_shot_job disposition=benign allowed=true reason=benign
|
||||
time=2026-10-04T08:45:31.852+02:00 level=INFO msg="restore-test: scratch guest torn down" vmid=990000
|
||||
--- restore-test record ---
|
||||
{
|
||||
"source_archive": "local:backup/vzdump-lxc-9201-2026_10_04-04_34_55.tar.zst",
|
||||
"source_tier": "local",
|
||||
"scratch_vmid": 990000,
|
||||
"pass": true,
|
||||
"verified": "boot+running",
|
||||
"tested_at": "2026-10-04T06:45:31Z",
|
||||
"duration_seconds": 79.434773547,
|
||||
"mount_parity": "ok",
|
||||
"mount_inventory": [
|
||||
"mp0=/var/lib/felhom (70G)",
|
||||
"mp8=/mnt/felhom-drives (throwaway for the archived bind)",
|
||||
"mp9=/etc/felhom-bootstrap (throwaway for the archived bind)"
|
||||
]
|
||||
}
|
||||
space preflight passed: storage=nvme-scratch required=24520159232 avail=876487208960
|
||||
=== selftest=restore-test OK (scratch 990000 restored+booted+verified+torn-down in 1m19s) ===
|
||||
@@ -0,0 +1,10 @@
|
||||
# R-833 hub red-proofs, 2026-10-04 (hub tree before v0.129.0 release). Each mutation applied, test run, reverted.
|
||||
RP1 raise ignored (OpenWindowFor keeps the default cap):
|
||||
--- FAIL: TestWindow_RaisedCapIsOneWindowOnly — raised grant: {Granted:true WindowID:2 ... MaxRemove:20} — want MaxRemove 30
|
||||
RP2 grant not consumed (TakeOffsiteWindowGrant does not clear the setting):
|
||||
--- FAIL: TestWindow_RaisedCapIsOneWindowOnly — the raised grant was not consumed
|
||||
RP3 close check uses the default cap instead of the window's own:
|
||||
--- FAIL: TestWindow_RaisedCapIsOneWindowOnly — a drop of 28 under a raised cap of 30 alarmed — the close check ignored the window's own cap
|
||||
RP4 the box API window-open sets a grant from a max_remove in its body:
|
||||
--- FAIL: TestOffsiteWindow_BoxCannotGrantItself — a box call left a grant (ok=true max=400)
|
||||
After revert: ok internal/offsitekeys, ok internal/api
|
||||
@@ -0,0 +1,20 @@
|
||||
# R-833 lab proof — 2026-10-04 (a lab repo on DooPlex scratch, not a demo box's)
|
||||
|
||||
restic 0.14.0 from the controller image `felhom-controller:0.290.0` (`docker run --rm --entrypoint sh`), a local repo,
|
||||
password `lab-only-password` (a lab value). 98 daily snapshots, `--host demo-lab --tag felhom-offsite`, dated 97..0 days
|
||||
back — the shape of a box whose windows were off for three months.
|
||||
|
||||
The box's own guard (`offsiteGuard`, controller v0.290.0 source) was run on the repo's REAL `snapshots --json` and the
|
||||
policy's REAL `forget --dry-run --json` by `offbox_window_lab_test.go.txt` (copy it into
|
||||
`controller/internal/backup/` and set `OFFSITE_LAB_DIR` to rerun; it skips without it, so it is not committed).
|
||||
|
||||
| Step | Result |
|
||||
|---|---|
|
||||
| Honest plan | 98 snapshots, the policy removes **85** |
|
||||
| Default cap (hub `MaxRemove` = half) | **49 → REFUSED**: `the plan would remove 85 snapshots, more than one week's retention may (49)` |
|
||||
| Operator-raised cap 90 (one window) | **85 ids, no refusal** |
|
||||
| `restic forget <those 85 ids> --prune` | rc 0; **98 → 13**; `restic check` rc 0 |
|
||||
| Next window, default cap again (6) | plan 0, **no refusal** — the cap only needed raising once |
|
||||
|
||||
The hub half (the grant is one-shot, the next window has the default cap, the close check uses the window's own cap, a
|
||||
box cannot grant itself) is in `hub-redproofs.txt` and the hub suite.
|
||||
@@ -0,0 +1,64 @@
|
||||
package backup
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// R-833 LAB (not run in CI: needs OFFSITE_LAB_DIR from the lab script). Reads a REAL restic 0.14.0
|
||||
// repo's `snapshots --json` and the policy's `forget --dry-run --json`, and runs the box's own guard
|
||||
// with the hub's default cap and with an operator-raised cap. Writes the ids the guard would remove to
|
||||
// $OFFSITE_LAB_DIR/ids.txt for the lab script to prune. Evidence: audits/backup-close-2026-10-04/partB/.
|
||||
func TestOffsiteGuard_R833_LabRepo(t *testing.T) {
|
||||
dir := os.Getenv("OFFSITE_LAB_DIR")
|
||||
if dir == "" {
|
||||
t.Skip("lab only: set OFFSITE_LAB_DIR")
|
||||
}
|
||||
var all []guardSnap
|
||||
b, err := os.ReadFile(filepath.Join(dir, "snaps.json"))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := json.Unmarshal(b, &all); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
b, err = os.ReadFile(filepath.Join(dir, "plan.json"))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
var groups []struct {
|
||||
Remove []guardSnap `json:"remove"`
|
||||
}
|
||||
if err := json.Unmarshal(b, &groups); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
var plan []guardSnap
|
||||
for _, g := range groups {
|
||||
plan = append(plan, g.Remove...)
|
||||
}
|
||||
now := time.Now()
|
||||
def := len(all) / 2 // the hub's MaxRemove(count): half, minimum 5
|
||||
if def < 5 {
|
||||
def = 5
|
||||
}
|
||||
_, refuse := offsiteGuard(all, plan, now, now, def)
|
||||
t.Logf("snapshots=%d plan=%d default cap=%d → refusal: %q", len(all), len(plan), def, refuse)
|
||||
raised := 0
|
||||
if v := os.Getenv("OFFSITE_LAB_RAISED"); v != "" {
|
||||
for _, c := range v {
|
||||
raised = raised*10 + int(c-'0')
|
||||
}
|
||||
}
|
||||
if raised == 0 {
|
||||
return
|
||||
}
|
||||
ids, refuse2 := offsiteGuard(all, plan, now, now, raised)
|
||||
t.Logf("raised cap=%d → %d id(s), refusal: %q", raised, len(ids), refuse2)
|
||||
if err := os.WriteFile(filepath.Join(dir, "ids.txt"), []byte(strings.Join(ids, "\n")), 0o600); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
@@ -73,6 +73,12 @@ pct list | awk '{print $1}' | grep -x -e <TARGET-VMID> -e <SOURCE-VMID>
|
||||
|
||||
## 3. Restore, then fix the binds BEFORE first boot
|
||||
|
||||
> **Restoring BESIDE a live original** (a copy to read, a check, a drill)? **Do not use this section.** Use
|
||||
> `felhom.eu/scripts/felhom-restore-beside.sh <scratch-vmid> <volid> <storage>`: onboot 0 at restore, every host-path
|
||||
> bind removed, every NIC down, never started, config read back (R-834, proven live 2026-10-04). This section is for
|
||||
> a **replaced host**, where the original is gone and the binds are right. A bare `pct restore` keeps the archive's
|
||||
> `onboot: 1`, so a host reboot would start it.
|
||||
|
||||
Restore **without starting** the guest. `pct restore` does not auto-start, but never pass anything that
|
||||
would, and do not `pct start` until §4 passes.
|
||||
|
||||
|
||||
@@ -45,14 +45,16 @@ recovered with the household's recovery code — the same as restoring from ep0)
|
||||
`namespace <customer>` / `username root@pam!<name>`.
|
||||
**Do not use `pvesm add pbs` without `--password`:** it validates with the password from its command line, fails 401,
|
||||
and on failure DELETES the `.pw`/`.enc` files you placed (measured). Passing `--password` puts the token on argv.
|
||||
3. `pvesm list <id>` → the household's snapshots (measured: 2 s). `pct restore <scratch VMID> <id>:backup/ct/<vmid>/<time>
|
||||
--storage <dir storage> --unique 1` (measured: **186 s for a 15 GB-logical / 14 GB-on-disk backup** over the LAN, key
|
||||
fingerprint printed by the restore).
|
||||
4. **⚠ BEFORE ANYTHING ELSE — the restored config is the PRODUCTION one:** `onboot: 1`, `mp8` bound to the host's REAL
|
||||
household drives (`/mnt/felhom-drives`) and `mp9` to the original guest's bootstrap. Starting it, or a host reboot,
|
||||
runs a second controller for the same household against the same drives. On a restore BESIDE the original:
|
||||
`pct set <vmid> --onboot 0 --delete mp8,mp9` immediately (R-834). On a true replacement host, where the original is
|
||||
gone, the binds are what you want.
|
||||
3. `pvesm list <id>` → the household's snapshots (measured: 2 s). Then restore with the safe script (R-834):
|
||||
`felhom-restore-beside.sh <scratch VMID> <id>:backup/ct/<vmid>/<time> <dir storage>` (copy it from
|
||||
`felhom.eu/scripts/`; measured 2026-10-03 by hand: **186 s for a 15 GB-logical / 14 GB-on-disk backup** over the LAN).
|
||||
4. **⚠ Why the script, not a bare `pct restore`:** the archive's config is the PRODUCTION one — `onboot: 1`, `mp8` bound
|
||||
to the host's REAL household drives (`/mnt/felhom-drives`) and `mp9` to the original guest's bootstrap. Started, or
|
||||
after a host reboot, it is a second controller for the same household on the same drives. The script restores with
|
||||
`--onboot 0`, removes every host-path bind, takes every NIC down, never starts it, and reads the config back
|
||||
(proven live 2026-10-04, `audits/backup-close-2026-10-04/partA/`). On a true replacement host, where the original is
|
||||
gone, the binds are what you want: use `RUNBOOK-manual-guest-restore.md` §3 or the agent's DR bring-up instead (the
|
||||
DR bring-up refuses beside a live original since agent v0.139.0).
|
||||
5. Read the data without starting it: `pct mount <vmid>` → `/var/lib/lxc/<vmid>/rootfs` (measured: 1 s; rootfs, the
|
||||
controller data volume and `/var/lib/felhom` present, `settings.json` dated 10 min before the backup) → `pct unmount`.
|
||||
6. Teardown: `pct destroy <vmid> --purge`; `pvesm remove <id>` (removes the entry and its priv files — never the
|
||||
|
||||
Reference in New Issue
Block a user