Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The restore test off, demo-hp 9201 repaired, agent v0.133.0 + hub v0.124.0, controller v0.270.0 (2026-09-24 evening)
Brief: "the restore test switched off on the demo boxes, the HP customer box repaired, the host agent fixed so a
restore test can never fill a box's disk again; then four controller leftovers". Architecture read: 03-host-agent.md
§8/§10, 07, 08 §6.2, 09 §3/§6.4, runbooks/target-selection.md.
Not done, or changed
- Agent v0.133.0 is released, NOT delivered. It reaches a box only with the operator's signed
agent_updatejob (R-530). Until then the scheduled restore test stays OFF on both demo hosts. - Part A used
-1, not the brief's0.restore_test_eval_interval_seconds: 0means "the 6-hour default" (config.goRestoreTestEvalInterval); only a negative value disables. The start-up line says "restore-test cadence disabled" on both hosts. With the cadence off no evaluation runs, so there is no "nothing due" line to quote. - Part C rule 1 changed: "archive size × 1.2 + 5 GiB" would NOT have prevented the incident. The size is the UNCOMPRESSED one (vzdump log / PBS snapshot size).
- Part C rule 4 needed a hub release (v0.124.0). The existing
storage_fill_criticalfired only at 95 %, on data only, per household per hour, and only at the next 15-minute report. Now: a thin pool is critical at 90 % of data or metadata, one alarm per pool per 6 hours, and the agent asks for a report the moment a pool crosses 90 %. - Live case (b) refused and case (c) did not run. On demo-hp, 9201's restore needs 30.3 GiB and the pool has 22.1 GiB free. So no full restore test fits under the 80 % limit. The forced clean-up failure needs a scratch guest, which cannot be created there. Both are covered by unit tests only.
- Part B found damage fsck could not see: the Redis append-only files of docmost and romm were cut off by the full
pool. With the operator's yes, each Redis folder was copied aside and cut at the last complete write with
redis-check-aof --fix(2,943 B / 6,631 B dropped). Meanwhile the box had stopped both apps by itself (decision 28). - R-674 has no live proof: R-679 now refuses the only case that reached that log line.
- R-679 removed a "repair path" that a test comment named: a same-version Update. Restart is the repair path.
Interventions: 0 on product behaviour. The Redis repair was an operator-approved data repair. One agent release, one hub release, one controller release.
Claims in the brief that turned out wrong (or right)
- An eval interval of 0 disables only the schedule and leaves the on-demand test working — half wrong. 0 is
the 6-hour default; negative disables. The on-demand path (
--selftest=restore-test) does not read the cadence — TRUE, and used live. pct fsckcan check the/var/lib/felhommount by--device— TRUE (--device mp0).- demo-hp has a second eligible storage for the restore test — WRONG.
nvme-scratchtakesrootdir, but the agent holds only the inheritedDatastore.Auditthere, notDatastore.AllocateSpace. - R-673's lock came from the pool-full event — WRONG. The 06:59 and 07:42 backups failed with "No space left on
device" on
local, the host ROOT disk (R-684). The lock is from the 07:42 clean-up. The pool filled at 10:35. - "Archive size × 1.2 + 5 GiB" is enough — WRONG. 9201: a 6.9 GB file, a 22.6 GB restore.
- A pool nearly full is only a log line — partly wrong. The hub DID mail
storage_fill_criticalat 100 %, but late and at the generic bands.
Part A — the scheduled restore test off (A-restore-test-off.txt)
Both hosts: only backup.restore_test_eval_interval_seconds changed (saved copy /etc/felhom-agent/agent.json.pre-r672,
verified equal apart from that key). Start-up: "backup: restore-test cadence disabled". Peti's box untouched: it gets
the fix only through a signed agent.
Part B — 9201 repaired (B1…B9)
Pool 58.9 % data / 2.65 % metadata. rootfs = vm-9201-disk-0, /var/lib/felhom = mp0 (vm-9201-disk-1). Stop 4.8 s.
pct fsck: rootfs replayed its journal; mp0 fixed 17 "deleted inode has zero dtime" and one orphan block; second
pass clean on both (rc 0). Start; the controller took the 0.269.1 floor by itself; both disks writable; 21 containers
up; hub /hosts: demo-hp ONLINE. Then the Redis repair above; docmost and romm started through the product (Start
lifted the box's hold), 0 restarts.
Part C — agent v0.133.0 (tag 9bdb4da, sha256 3aa30345…e69b6, verified by an anonymous download) + hub v0.124.0
Red-proofs (redproofs/C-*): no preflight → the 2026-09-24 restore issued again; file size used → "the compressed
file size was used"; no timer → "the leaked scratch was not destroyed by the timer"; sweep without the gate → "the
sweep ran while a backup held the gate"; no 90 % edge → "0 report requests, want 1"; skips dropped → "a space refusal
never reached the host report"; hub generic bands → a 91 % pool only warns, a metadata-full pool raises nothing; hub
old key → "a second pool filling in the same hour was silenced".
Live on demo-hp (on-demand self-test, cadence off, pool 58.99 % before and after, nothing created): (a) factor 10 →
"needs 215.5 GiB free, has 22.1 GiB", exit 4; (b) normal → "restoring 21.1 GiB (vzdump log: total bytes written)
needs 30.3 GiB free, has 22.1 GiB", exit 4.
R-673: the stale-lock sweep now runs every 10 minutes, under the one-heavy-operation gate.
Part D — controller v0.270.0, floor 0.270.0 (D/)
R-679 (409 already_current, hu + en, live), R-681 (interrupted install reported and cleaned, live), R-669 (the unit
keeps the pinned health check; a restore resets the applied record; live), R-674 (unit only). Both demo boxes
arrived on 0.270.0 in ~12 s.
Register
Before this session 344 rows / 685,662 B; after 341 rows / 685,148 B. Opened R-684; closed R-669, R-674, R-679, R-681; R-672 and R-673 updated to "fixed in v0.133.0, awaiting delivery".
Teardown — three layers
- Machine: 9202 back on the live catalog with its saved config, on controller 0.270.0 (the floor); the throwaway apps (actualbudget, mealie, wishlist) removed through the product — no containers, volumes or markers left. 9201 (demo-hp) repaired, running, on 0.270.0.
- Host: demo-hp and demo-felhom: only the agent config key changed (saved copies beside them); the live-test
binary and its test config deleted from demo-hp
/tmp;pct list9201 + 9202, nothing created. - Hub: v0.124.0 deployed; floor 0.270.0 (MinAgent 0.131.0); drill repo reset to the live catalog.
To turn the restore test back on (after the signed agent arrives on a box)
On that host: cp /etc/felhom-agent/agent.json.pre-r672 /etc/felhom-agent/agent.json && systemctl restart felhom-agent.
The saved copy had no restore_test_eval_interval_seconds key (= the 6-hour default). Check the journal for "restore-test
scheduler starting". On demo-hp the test will then REFUSE 9201's restore for space, correctly, until the pool has room
(R-684 is about the separate backup storage).