Files
felhom-agent/REPORT.md
T
admin 567cf9f401 docs: REPORT + CONTEXT for S4.1 (tier-aware deadline + teardown phantom confirmed)
Live diagnosis on the agent-token path: offsite restore-test pass:true,
verified:boot+running, teardown succeeded (no 403), scratch band clean. The
VM.Allocate follow-up was a PHANTOM — the earlier 403 was the 10m-timeout firing
teardown against a not-yet-pool-associated guest; with the timeout fixed the
guest is in /pool/felhom and teardown is authorized. No ACL/host-install change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PSK5g6qYLknKj8u3QAFEr6
2026-07-04 19:54:46 +02:00

4.2 KiB

REPORT — S4.1: unattended offsite restore-test (tier-aware deadline + teardown diagnosis)

Date: 2026-07-04 · Class: implementation + live diagnosis on felhom-pve (agent token). Design of record: felhom.eu/documentation/architecture/06-offsite-connectivity.md S4 row; runbook offsite-endpoint.md §4b.

1. Baselines & commits (felhom-agent → main)

Start: fee1fcf (v0.67.0).

  • dc70e15 v0.68.0 — tier-aware restore-task deadline (Part 1).
  • this commit — REPORT + CONTEXT (the diagnosis outcome).

2. Part 1 — tier-aware restore-task deadline

The restore-task wait (restoretest.go ~L246) used WaitOptions{} → the 10m default. A WAN (pbs) restore of a large guest exceeds it, so the wait expired mid-restore and teardown then fired against a still-restoring guest. Fix:

  • RestoreTestSpec.RestoreTaskTimeout (0 → 10m default); the restore WaitTask passes it.
  • config.BackupConfig.RestoreTestPBSRestoreTimeoutSeconds + accessor RestoreTestPBSRestoreTimeout() (positive as-is, else 120m).
  • main.restoreTaskTimeout(cfg, tier) — the configured PBS timeout only when SourceTier=="pbs", else 0. Set at BOTH spec builds (scheduler + selftest). Local tier unchanged (10m).
  • Tests: WaitOptions.Timeout pbs→120m / local→0 (red-proofed against WaitOptions{}: pbs assertion fails); accessor contract (0/neg→120m, positive as-is). Full go test ./... green.

3. Part 2 — teardown diagnosis (the "grant VM.Allocate" follow-up) — PHANTOM CONFIRMED

Ran the restore-test on the agent-token path (selftest = the identical RunRestoreTest the scheduler runs; all PVE ops authenticate as felhom-agent@pve!agent), sourcing the offsite (pbs-tier) backup of guest 9201, with v0.68.0's 120m timeout.

Scenario B (done-criterion) — GREEN:

source_archive: felhom-offsite:backup/ct/9201/2026-07-04T14:55:44Z
source_tier: pbs · pass: true · verified: boot+running · scratch_vmid: 990000 · duration 4m49s
restore → net-link-down → boot (running) → teardown → scratch band clean (no leak)

Scenario C — the teardown succeeded as the agent token; (b) is a PHANTOM:

gate decision: class=guest_destroy guest=990000 allowed=true reason=benign
restore-test: scratch guest torn down vmid=990000
pct list | grep 9900 → no scratch guests

No 403, no VM.Allocate error. Root cause of the earlier 403 (REPORT of S4): the 10m wait expired mid-restore, so teardown hit a guest that existed but was not yet pool-associated (the pool membership lands with restore completion) — a bare /vms/990000 the pool ACL didn't cover. With the timeout fixed → the restore completes → the guest is a member of /pool/felhom (restored with Pool=DefaultPool, restoretest.go L229) → its destroy is authorized by the existing FelhomAgentGuest grant (VM.Allocate + Pool.Allocate at /pool/felhom). No ACL change, no host-install edit — the pool-scoped containment (SPIKE-pool-scoped-acl-2026-07-01) stands.

(Note: this run's restore was fast — home downloadupload; the earlier ~10-min stall was link contention. The 120m ceiling covers the slow case regardless — for an unattended nightly test a false timeout is the failure mode to prevent.)

4. Live state after / cleanup

Agent v0.68.0 live on felhom-pve; the temporary local_backup_target=felhom-offsite (set so the selftest sourced the pbs-tier backup) reverted to local; WARN still 0; scratch band clean; temp logs removed. Rollback binary: felhom-agent.bak-0.67.0.

5. NOT yet live-validated / follow-ups / observations

  • The restore-test is now unattended-capable for the pbs tier (deadline fixed; teardown proven as the token). The scheduler cadence stays disabled on the demo (restore_test_cadence_seconds=0) — the proof used the identical selftest path.
  • Tier-1(local, frequent)/Tier-2(offsite, periodic) split still needed before offsite can be the DEFAULT backup target (the controller's ~30-min cadence + a single target = near-continuous offsite uploads). Unchanged from S4.
  • Deferred: S5 DR consume (WG key + offsite restore with only R); hub-driven tenancy automation; true-CGNAT-SIM retest. OPERATOR: publish 0.68.0 + Day-0 vouch.