Files
felhom.eu/STATUS.md
T

3.2 KiB
Raw Blame History

STATUS — what works, what's broken, what's next

Updated 2026-09-24 (evening). The demo boxes are on controller 0.270.0. The HP customer box is repaired. A fixed host agent (0.133.0) is released and waits for your signature. Until then the automatic restore test is off on both demo boxes.

Decisions I took on my own (you may reverse them).

  • The space check uses the real, uncompressed size, not the backup file size. The rule in your brief ("file size × 1.2 + 5 GB") would NOT have stopped last night's accident. The HP box's backup file is 7 GB, but restoring it writes 22.6 GB.
  • To switch the test off I used the value −1, not 0. In this setting, 0 means "every 6 hours". Only a negative number turns it off.
  • The hub got a small release too (0.124.0). A thin disk pool now counts as urgent at 90 % (data or its bookkeeping part), with one alarm per pool per 6 hours. Before, the alarm came only at 95 %, up to 15 minutes late, and one full pool could silence another.

What I did, and it worked.

  • The HP customer box is repaired. I stopped it, checked both disks (small damage fixed, second check clean), and started it. All apps came back, it took the current version by itself, and the hub shows it online. Two apps' cache files (Redis) had been cut off by the full disk; with your yes I kept copies and repaired them. Both apps run.
  • The box's own crash-loop stop worked for real: while the cache was broken, the box stopped those two apps and told the hub, by itself.
  • The new agent never starts a restore test that does not fit. Tested on the HP box: it refused, said why, and created nothing.
  • Four controller fixes (0.270.0): Update on an app that is already current is refused, with no restart. An install cut off by a restart is now cleaned up, reported and shown on the page. After a restore, the box keeps the right health check. One log line is corrected.

What broke, or is not done.

  • The HP box's own backups have been failing every day since yesterday. Its host disk has 4 GB free, but one backup needs 7 GB, and three old ones are kept. Nothing fixes this by itself.
  • No full restore test fits on the HP box today (it needs 30 GB free, 22 GB is free). The new agent will refuse it every time, correctly. So a passing test on the HP box is not proven yet.

Rows. 1 opened, 4 closed, 2 updated. The list went from 344 to 341.

What needs you.

  1. Sign the agent update (0.133.0) for demo-hp and demo-felhom. Version 0.133.0, sha256 3aa303452b8c6be58573d00af01a0ab4a0d97e4f885ffecd0144a18ac24e69b6, one felhom-opsign -op agent_update job per box, the same way as 0.131.0. If you do nothing, the boxes keep the old agent and the automatic restore test stays off.
  2. After the new agent is on a box, turn the restore test back on. On that host, remove the key restore_test_eval_interval_seconds from backup in /etc/felhom-agent/agent.json (the saved copy agent.json.pre-r672 has it absent), then systemctl restart felhom-agent. If you do nothing, no backup is proven restorable on that box.
  3. Decide what to do with the HP box's backup storage: keep fewer old backups, back up somewhere else, or add disk. If you do nothing, its whole-box backup keeps failing every night.