d53bd08442
gates / gates (push) Successful in 28s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
90 lines
7.0 KiB
Markdown
90 lines
7.0 KiB
Markdown
# The restore test off, demo-hp 9201 repaired, agent v0.133.0 + hub v0.124.0, controller v0.270.0 (2026-09-24 evening)
|
||
|
||
Brief: "the restore test switched off on the demo boxes, the HP customer box repaired, the host agent fixed so a
|
||
restore test can never fill a box's disk again; then four controller leftovers". Architecture read: `03-host-agent.md`
|
||
§8/§10, `07`, `08` §6.2, `09` §3/§6.4, `runbooks/target-selection.md`.
|
||
|
||
## Not done, or changed
|
||
|
||
- **Agent v0.133.0 is released, NOT delivered.** It reaches a box only with the operator's signed `agent_update` job
|
||
(R-530). Until then the scheduled restore test stays OFF on both demo hosts.
|
||
- **Part A used `-1`, not the brief's `0`.** `restore_test_eval_interval_seconds: 0` means "the 6-hour default"
|
||
(`config.go` `RestoreTestEvalInterval`); only a negative value disables. The start-up line says "restore-test cadence
|
||
disabled" on both hosts. With the cadence off no evaluation runs, so there is no "nothing due" line to quote.
|
||
- **Part C rule 1 changed:** "archive size × 1.2 + 5 GiB" would NOT have prevented the incident. The size is the
|
||
UNCOMPRESSED one (vzdump log / PBS snapshot size).
|
||
- **Part C rule 4 needed a hub release (v0.124.0).** The existing `storage_fill_critical` fired only at 95 %, on data
|
||
only, per household per hour, and only at the next 15-minute report. Now: a thin pool is critical at 90 % of data
|
||
or metadata, one alarm per pool per 6 hours, and the agent asks for a report the moment a pool crosses 90 %.
|
||
- **Live case (b) refused and case (c) did not run.** On demo-hp, 9201's restore needs 30.3 GiB and the pool has 22.1
|
||
GiB free. So no full restore test fits under the 80 % limit. The forced clean-up failure needs a scratch guest,
|
||
which cannot be created there. Both are covered by unit tests only.
|
||
- **Part B found damage fsck could not see:** the Redis append-only files of docmost and romm were cut off by the full
|
||
pool. With the operator's yes, each Redis folder was copied aside and cut at the last complete write with
|
||
`redis-check-aof --fix` (2,943 B / 6,631 B dropped). Meanwhile the box had stopped both apps by itself (decision 28).
|
||
- **R-674 has no live proof:** R-679 now refuses the only case that reached that log line.
|
||
- **R-679 removed a "repair path"** that a test comment named: a same-version Update. Restart is the repair path.
|
||
|
||
**Interventions: 0** on product behaviour. The Redis repair was an operator-approved data repair. **One agent
|
||
release, one hub release, one controller release.**
|
||
|
||
## Claims in the brief that turned out wrong (or right)
|
||
|
||
1. *An eval interval of 0 disables only the schedule and leaves the on-demand test working* — **half wrong.** 0 is
|
||
the 6-hour default; negative disables. The on-demand path (`--selftest=restore-test`) does not read the cadence —
|
||
TRUE, and used live.
|
||
2. *`pct fsck` can check the `/var/lib/felhom` mount by `--device`* — **TRUE** (`--device mp0`).
|
||
3. *demo-hp has a second eligible storage for the restore test* — **WRONG.** `nvme-scratch` takes `rootdir`, but the
|
||
agent holds only the inherited `Datastore.Audit` there, not `Datastore.AllocateSpace`.
|
||
4. *R-673's lock came from the pool-full event* — **WRONG.** The 06:59 and 07:42 backups failed with "No space left on
|
||
device" on `local`, the host ROOT disk (R-684). The lock is from the 07:42 clean-up. The pool filled at 10:35.
|
||
5. *"Archive size × 1.2 + 5 GiB" is enough* — **WRONG.** 9201: a 6.9 GB file, a 22.6 GB restore.
|
||
6. *A pool nearly full is only a log line* — **partly wrong.** The hub DID mail `storage_fill_critical` at 100 %, but
|
||
late and at the generic bands.
|
||
|
||
## Part A — the scheduled restore test off (`A-restore-test-off.txt`)
|
||
Both hosts: only `backup.restore_test_eval_interval_seconds` changed (saved copy `/etc/felhom-agent/agent.json.pre-r672`,
|
||
verified equal apart from that key). Start-up: "backup: restore-test cadence disabled". Peti's box untouched: it gets
|
||
the fix only through a signed agent.
|
||
|
||
## Part B — 9201 repaired (`B1…B9`)
|
||
Pool 58.9 % data / 2.65 % metadata. rootfs = `vm-9201-disk-0`, `/var/lib/felhom` = `mp0` (`vm-9201-disk-1`). Stop 4.8 s.
|
||
`pct fsck`: rootfs replayed its journal; mp0 fixed 17 "deleted inode has zero dtime" and one orphan block; second
|
||
pass clean on both (rc 0). Start; the controller took the 0.269.1 floor by itself; both disks writable; 21 containers
|
||
up; hub `/hosts`: demo-hp ONLINE. Then the Redis repair above; docmost and romm started through the product (Start
|
||
lifted the box's hold), 0 restarts.
|
||
|
||
## Part C — agent v0.133.0 (tag `9bdb4da`, sha256 `3aa30345…e69b6`, verified by an anonymous download) + hub v0.124.0
|
||
Red-proofs (`redproofs/C-*`): no preflight → the 2026-09-24 restore issued again; file size used → "the compressed
|
||
file size was used"; no timer → "the leaked scratch was not destroyed by the timer"; sweep without the gate → "the
|
||
sweep ran while a backup held the gate"; no 90 % edge → "0 report requests, want 1"; skips dropped → "a space refusal
|
||
never reached the host report"; hub generic bands → a 91 % pool only warns, a metadata-full pool raises nothing; hub
|
||
old key → "a second pool filling in the same hour was silenced".
|
||
Live on demo-hp (on-demand self-test, cadence off, pool 58.99 % before and after, nothing created): (a) factor 10 →
|
||
"needs 215.5 GiB free, has 22.1 GiB", exit 4; (b) normal → "restoring 21.1 GiB (vzdump log: total bytes written)
|
||
needs 30.3 GiB free, has 22.1 GiB", exit 4.
|
||
R-673: the stale-lock sweep now runs every 10 minutes, under the one-heavy-operation gate.
|
||
|
||
## Part D — controller v0.270.0, floor 0.270.0 (`D/`)
|
||
R-679 (409 `already_current`, hu + en, live), R-681 (interrupted install reported and cleaned, live), R-669 (the unit
|
||
keeps the pinned health check; a restore resets the applied record; live), R-674 (unit only). Both demo boxes
|
||
arrived on 0.270.0 in ~12 s.
|
||
|
||
## Register
|
||
Before this session **344 rows / 685,662 B**; after **341 rows / 685,148 B**. Opened R-684; closed R-669, R-674,
|
||
R-679, R-681; R-672 and R-673 updated to "fixed in v0.133.0, awaiting delivery".
|
||
|
||
## Teardown — three layers
|
||
- **Machine:** 9202 back on the live catalog with its saved config, on controller 0.270.0 (the floor); the throwaway
|
||
apps (actualbudget, mealie, wishlist) removed through the product — no containers, volumes or markers left.
|
||
9201 (demo-hp) repaired, running, on 0.270.0.
|
||
- **Host:** demo-hp and demo-felhom: only the agent config key changed (saved copies beside them); the live-test
|
||
binary and its test config deleted from demo-hp `/tmp`; `pct list` 9201 + 9202, nothing created.
|
||
- **Hub:** v0.124.0 deployed; floor 0.270.0 (MinAgent 0.131.0); drill repo reset to the live catalog.
|
||
|
||
## To turn the restore test back on (after the signed agent arrives on a box)
|
||
On that host: `cp /etc/felhom-agent/agent.json.pre-r672 /etc/felhom-agent/agent.json && systemctl restart felhom-agent`.
|
||
The saved copy had no `restore_test_eval_interval_seconds` key (= the 6-hour default). Check the journal for "restore-test
|
||
scheduler starting". On demo-hp the test will then REFUSE 9201's restore for space, correctly, until the pool has room
|
||
(R-684 is about the separate backup storage).
|